Websites blocking Google-Extended
A robots.txt control token rather than a separate crawler. Disallowing Google-Extended tells Google not to use a site's content for Gemini model training and grounding, without affecting Google Search indexing.
Operated by Google · Training crawler · operator documentation
* block
2.63%
restrict it by name
655K domains name it and close some of its paths
9.48%
name it in robots.txt
2.36 million domains give it rules of their own
Data for Aug 2026. All four percentages
divide by the same 24.91 million domains. Blocked counts both
explicit User-agent: Google-Extended blocks and blanket
* blocks the crawler inherits.
Historical mentions, HTTP Archive
High-authority sites blocking Google-Extended by name
| Domain | OPROpen Page Rank | Blocking since |
|---|---|---|
| linkedin.com | 9.82/10 | Aug 2026 |
| fbcdn.net | 9.40/10 | Aug 2026 |
| nytimes.com | 9.28/10 | Aug 2026 |
| tumblr.com | 9.27/10 | Aug 2026 |
| wiley.com | 9.14/10 | Aug 2026 |
| cookielaw.org | 9.10/10 | Aug 2026 |
| sagepub.com | 9.07/10 | Aug 2026 |
| bandcamp.com | 9.06/10 | Aug 2026 |
| tripadvisor.com | 9.05/10 | Aug 2026 |
| lnkd.in | 9.05/10 | Aug 2026 |
| congress.gov | 9.05/10 | Aug 2026 |
| bizjournals.com | 9.04/10 | Aug 2026 |
| msn.com | 9.04/10 | Aug 2026 |
| mayoclinic.org | 9.04/10 | Aug 2026 |
| pexels.com | 9.03/10 | Aug 2026 |
| theatlantic.com | 9.01/10 | Aug 2026 |
| hugedomains.com | 9.00/10 | Aug 2026 |
| uk.com | 9.00/10 | Aug 2026 |
| patreon.com | 8.99/10 | Aug 2026 |
| science.org | 8.95/10 | Aug 2026 |
| lefigaro.fr | 8.95/10 | Aug 2026 |
| usgs.gov | 8.95/10 | Aug 2026 |
| utexas.edu | 8.92/10 | Aug 2026 |
| huffpost.com | 8.91/10 | Aug 2026 |
| newyorker.com | 8.90/10 | Aug 2026 |
| claude.ai | 8.87/10 | Aug 2026 |
| francetvinfo.fr | 8.87/10 | Aug 2026 |
| jsfiddle.net | 8.86/10 | Aug 2026 |
| nami.org | 8.86/10 | Aug 2026 |
| smh.com.au | 8.86/10 | Aug 2026 |
Sites whose robots.txt names Google-Extended (or a legacy
alias) with a full block, ranked by Open Page Rank. Sites blocking every crawler with a blanket
* rule are counted above but not listed here. Since dates start at our first
observation of the rule.
Restricting some paths
- gravatar.com
- vimeo.com
- statista.com
- bbc.com
- hbr.org
- squarespace.com
- usatoday.com
- frontiersin.org
- iubenda.com
- theverge.com
- nbcnews.com
- investopedia.com
Crawler first seen in the wild around Sep 2023. How to block it and what each state means is on the methodology page.