Who blocks AI crawlers
How websites respond to AI crawlers in robots.txt, measured by our own monthly crawl of 44.1 million domains. Data for Aug 2026.
*
889K of 20.95 million with robots.txt
Percentages divide by the 24.91 million domains whose robots.txt policy we could determine this month, except where a tile names a different base. Full definitions in the methodology.
Every AI crawler we track
Blocked means the domain's robots.txt shuts the crawler out entirely, either by naming it
or via a blanket * rule. Restricted by name means the file names the crawler and
closes some of its paths. Named means the crawler appears in the file with its own rules,
whatever they say.
Training crawlers
| Crawler | Operator | Blocked | Restricted by name | NamedNamed in robots.txt |
|---|---|---|---|---|
| CCBot | Common Crawl | 10.01% (2.49M) | 624K | 2.37M |
| GPTBot | OpenAI | 9.99% (2.49M) | 690K | 2.55M |
| Bytespider | ByteDance | 9.92% (2.47M) | 603K | 2.29M |
| ClaudeBot | Anthropic | 9.74% (2.43M) | 668K | 2.47M |
| Google-Extended | 9.51% (2.37M) | 655K | 2.36M | |
| Applebot-Extended | Apple | 9.44% (2.35M) | 627K | 2.23M |
| meta-externalagent | Meta | 9.36% (2.33M) | 626K | 2.21M |
| Webzio-Extended | Webz.io | 4.09% (1.02M) | 576K | 755K |
| cohere-training-data-crawler | Cohere | 4.05% (1.01M) | 592K | 815K |
| ImagesiftBot | Hive | 3.96% (987K) | 23,726 | 167K |
| FacebookBot | Meta | 3.93% (978K) | 593K | 741K |
| Diffbot | Diffbot | 3.93% (978K) | 29,939 | 171K |
| AI2Bot | Allen Institute for AI | 3.76% (938K) | 574K | 660K |
| img2dataset | Open source (rom1504) | 3.65% (910K) | 572K | 612K |
| GoogleOther | 3.64% (907K) | 587K | 656K | |
| PanguBot | Huawei | 3.63% (903K) | 1,461 | 32,679 |
| SBIntuitionsBot | SB Intuitions | 3.58% (892K) | 1,004 | 20,401 |
| Google-CloudVertexBot | 3.57% (888K) | 5,963 | 26,409 |
Search and AI assistants
| Crawler | Operator | Blocked | Restricted by name | NamedNamed in robots.txt |
|---|---|---|---|---|
| Amazonbot | Amazon | 9.87% (2.46M) | 614K | 2.29M |
| PetalBot | Huawei | 8.51% (2.12M) | 39,644 | 1.32M |
| PerplexityBot | Perplexity | 3.90% (972K) | 118K | 455K |
| YouBot | You.com | 3.89% (970K) | 584K | 734K |
| DuckAssistBot | DuckDuckGo | 3.66% (912K) | 568K | 641K |
| Applebot | Apple | 3.62% (902K) | 68,478 | 183K |
| OAI-SearchBot | OpenAI | 3.61% (899K) | 99,650 | 293K |
| Claude-SearchBot | Anthropic | 3.55% (884K) | 46,569 | 120K |
| Kagibot | Kagi | 3.51% (874K) | 1,547 | 2,698 |
User-request fetchers
| Crawler | Operator | Blocked | Restricted by name | NamedNamed in robots.txt |
|---|---|---|---|---|
| ChatGPT-User | OpenAI | 3.96% (987K) | 119K | 451K |
| meta-externalfetcher | Meta | 3.60% (896K) | 30,509 | 82,973 |
| GoogleAgent-Mariner | 3.57% (889K) | 823 | 17,482 | |
| Perplexity-User | Perplexity | 3.56% (887K) | 48,268 | 127K |
| MistralAI-User | Mistral AI | 3.55% (885K) | 12,335 | 34,768 |
| Claude-User | Anthropic | 3.54% (881K) | 55,444 | 130K |
Blocking by segment
The same measures cut four ways. High-authority and high-traffic sites block AI crawlers far more often than the long tail, and platform defaults decide the numbers for millions of sites at once.
By authority (Open Page Rank)
| Segment | Robots.txt known | GPTBot | ClaudeBot | Name any AI crawler | Block allBlock all crawlers |
|---|---|---|---|---|---|
| OPR 9.0 and up | 324 | 12.96% | 13.27% | 20.06% | 8.87% |
| OPR 8.0 to 8.9 | 6,521 | 15.83% | 15.11% | 19.03% | 8.37% |
| OPR 7.0 to 7.9 | 33,025 | 14.66% | 13.57% | 17.09% | 6.49% |
| OPR 6.0 to 6.9 | 55,479 | 12.82% | 11.97% | 15.62% | 4.97% |
| OPR 5.0 to 5.9 | 112K | 12.78% | 12.18% | 15.50% | 6.12% |
By traffic rank (Chrome UX Report)
| Segment | Robots.txt known | GPTBot | ClaudeBot | Name any AI crawler | Block allBlock all crawlers |
|---|---|---|---|---|---|
| Top 1K | 569 | 21.62% | 20.04% | 24.25% | 7.78% |
| Top 10K | 5,285 | 19.47% | 19.11% | 23.90% | 6.15% |
| Top 100K | 51,242 | 18.40% | 17.87% | 20.73% | 6.15% |
| Top 1M | 519K | 14.43% | 14.14% | 16.55% | 4.75% |
| Top 10M | 5.19M | 8.03% | 7.75% | 14.21% | 1.84% |
By primary audience country
| Segment | Robots.txt known | GPTBot | ClaudeBot | Name any AI crawler | Block allBlock all crawlers |
|---|---|---|---|---|---|
| United States | 1.45M | 6.50% | 6.13% | 25.48% | 1.58% |
| Japan | 705K | 2.08% | 1.70% | 6.15% | 0.61% |
| India | 517K | 8.16% | 8.01% | 12.90% | 2.05% |
| Brazil | 513K | 9.40% | 9.28% | 19.54% | 3.81% |
| Germany | 498K | 3.89% | 3.28% | 9.11% | 1.27% |
| United Kingdom | 371K | 6.99% | 6.56% | 20.79% | 1.30% |
| France | 369K | 4.53% | 4.36% | 12.44% | 1.17% |
| Russia | 266K | 3.84% | 3.44% | 4.09% | 2.29% |
| Canada | 249K | 7.82% | 7.41% | 24.78% | 1.66% |
| Indonesia | 235K | 23.91% | 23.47% | 24.43% | 1.74% |
| Spain | 231K | 6.03% | 5.91% | 10.69% | 2.00% |
| Poland | 214K | 6.32% | 6.17% | 9.48% | 0.99% |
| Australia | 199K | 7.82% | 7.45% | 21.55% | 1.65% |
| Italy | 199K | 5.65% | 5.53% | 11.99% | 1.58% |
| Netherlands | 165K | 5.90% | 5.57% | 10.86% | 1.44% |
| Türkiye | 149K | 9.82% | 9.77% | 15.59% | 0.85% |
| Argentina | 122K | 7.96% | 7.47% | 11.56% | 1.56% |
| Belgium | 107K | 5.85% | 5.37% | 15.05% | 1.46% |
| Mexico | 106K | 6.04% | 6.01% | 15.13% | 1.60% |
| Vietnam | 105K | 11.07% | 11.01% | 12.36% | 2.04% |
By platform
| Segment | Robots.txt known | GPTBot | ClaudeBot | Name any AI crawler | Block allBlock all crawlers |
|---|---|---|---|---|---|
| Cloudflare | 5.36M | 14.86% | 14.56% | 18.58% | 2.40% |
| WordPress | 4.94M | 5.99% | 5.80% | 7.21% | 0.62% |
| Shopify | 566K | 4.62% | 4.54% | 3.38% | 2.99% |
| Wix | 544K | 0.83% | 0.81% | 95.33% | 0.08% |
| Squarespace | 299K | 2.08% | 2.07% | 96.65% | 0.19% |
| Joomla | 145K | 4.20% | 4.03% | 6.97% | 1.00% |
| Webflow | 123K | 3.92% | 3.63% | 6.37% | 1.04% |
| Drupal | 103K | 6.74% | 6.42% | 8.15% | 1.38% |
| Magento | 43,672 | 19.58% | 19.07% | 28.01% | 1.58% |
| Ghost | 6,560 | 7.52% | 7.33% | 9.51% | 0.67% |
Robots.txt known counts the segment's domains whose robots.txt policy we could determine; blocked and named percentages divide by that number, and the block-all column divides by the segment's domains with a robots.txt. Authority tiers come from Open Page Rank scores, traffic tiers from the newest Chrome UX Report month, country from where a site's Chrome users are, and platform from our technology detections.
Beyond robots.txt
How to read these numbers
Of the 44.1 million domains in this month's crawl: 30.91 million resolved in DNS, 28.64 million answered our fetch, and 24.91 million gave a definitive robots.txt answer (a readable file or a clean 404, which means everything is allowed). Percentages on this page use that last number unless labeled otherwise.
A robots.txt rule is a published preference, not an enforcement mechanism. We record what sites declare; we cannot see which crawlers comply. Method details on the methodology page.