Who blocks AI crawlers

How websites respond to AI crawlers in robots.txt, measured by our own monthly crawl of 44.1 million domains. Data for Aug 2026.

9.99% block GPTBot 2.49 million domains 9.74% block ClaudeBot 2.43 million domains 15.70% name at least one AI crawler 3.91 million domains 4.24% block all crawlers with * 889K of 20.95 million with robots.txt

Percentages divide by the 24.91 million domains whose robots.txt policy we could determine this month, except where a tile names a different base. Full definitions in the methodology.

Every AI crawler we track

Blocked means the domain's robots.txt shuts the crawler out entirely, either by naming it or via a blanket * rule. Restricted by name means the file names the crawler and closes some of its paths. Named means the crawler appears in the file with its own rules, whatever they say.

Training crawlers

Crawler Operator Blocked Restricted by name NamedNamed in robots.txt
CCBot Common Crawl 10.01% (2.49M) 624K 2.37M
GPTBot OpenAI 9.99% (2.49M) 690K 2.55M
Bytespider ByteDance 9.92% (2.47M) 603K 2.29M
ClaudeBot Anthropic 9.74% (2.43M) 668K 2.47M
Google-Extended Google 9.51% (2.37M) 655K 2.36M
Applebot-Extended Apple 9.44% (2.35M) 627K 2.23M
meta-externalagent Meta 9.36% (2.33M) 626K 2.21M
Webzio-Extended Webz.io 4.09% (1.02M) 576K 755K
cohere-training-data-crawler Cohere 4.05% (1.01M) 592K 815K
ImagesiftBot Hive 3.96% (987K) 23,726 167K
FacebookBot Meta 3.93% (978K) 593K 741K
Diffbot Diffbot 3.93% (978K) 29,939 171K
AI2Bot Allen Institute for AI 3.76% (938K) 574K 660K
img2dataset Open source (rom1504) 3.65% (910K) 572K 612K
GoogleOther Google 3.64% (907K) 587K 656K
PanguBot Huawei 3.63% (903K) 1,461 32,679
SBIntuitionsBot SB Intuitions 3.58% (892K) 1,004 20,401
Google-CloudVertexBot Google 3.57% (888K) 5,963 26,409

Search and AI assistants

Crawler Operator Blocked Restricted by name NamedNamed in robots.txt
Amazonbot Amazon 9.87% (2.46M) 614K 2.29M
PetalBot Huawei 8.51% (2.12M) 39,644 1.32M
PerplexityBot Perplexity 3.90% (972K) 118K 455K
YouBot You.com 3.89% (970K) 584K 734K
DuckAssistBot DuckDuckGo 3.66% (912K) 568K 641K
Applebot Apple 3.62% (902K) 68,478 183K
OAI-SearchBot OpenAI 3.61% (899K) 99,650 293K
Claude-SearchBot Anthropic 3.55% (884K) 46,569 120K
Kagibot Kagi 3.51% (874K) 1,547 2,698

User-request fetchers

Crawler Operator Blocked Restricted by name NamedNamed in robots.txt
ChatGPT-User OpenAI 3.96% (987K) 119K 451K
meta-externalfetcher Meta 3.60% (896K) 30,509 82,973
GoogleAgent-Mariner Google 3.57% (889K) 823 17,482
Perplexity-User Perplexity 3.56% (887K) 48,268 127K
MistralAI-User Mistral AI 3.55% (885K) 12,335 34,768
Claude-User Anthropic 3.54% (881K) 55,444 130K

Blocking by segment

The same measures cut four ways. High-authority and high-traffic sites block AI crawlers far more often than the long tail, and platform defaults decide the numbers for millions of sites at once.

By authority (Open Page Rank)

Segment Robots.txt known GPTBot ClaudeBot Name any AI crawler Block allBlock all crawlers
OPR 9.0 and up 324 12.96% 13.27% 20.06% 8.87%
OPR 8.0 to 8.9 6,521 15.83% 15.11% 19.03% 8.37%
OPR 7.0 to 7.9 33,025 14.66% 13.57% 17.09% 6.49%
OPR 6.0 to 6.9 55,479 12.82% 11.97% 15.62% 4.97%
OPR 5.0 to 5.9 112K 12.78% 12.18% 15.50% 6.12%

By traffic rank (Chrome UX Report)

Segment Robots.txt known GPTBot ClaudeBot Name any AI crawler Block allBlock all crawlers
Top 1K 569 21.62% 20.04% 24.25% 7.78%
Top 10K 5,285 19.47% 19.11% 23.90% 6.15%
Top 100K 51,242 18.40% 17.87% 20.73% 6.15%
Top 1M 519K 14.43% 14.14% 16.55% 4.75%
Top 10M 5.19M 8.03% 7.75% 14.21% 1.84%

By primary audience country

Segment Robots.txt known GPTBot ClaudeBot Name any AI crawler Block allBlock all crawlers
United States 1.45M 6.50% 6.13% 25.48% 1.58%
Japan 705K 2.08% 1.70% 6.15% 0.61%
India 517K 8.16% 8.01% 12.90% 2.05%
Brazil 513K 9.40% 9.28% 19.54% 3.81%
Germany 498K 3.89% 3.28% 9.11% 1.27%
United Kingdom 371K 6.99% 6.56% 20.79% 1.30%
France 369K 4.53% 4.36% 12.44% 1.17%
Russia 266K 3.84% 3.44% 4.09% 2.29%
Canada 249K 7.82% 7.41% 24.78% 1.66%
Indonesia 235K 23.91% 23.47% 24.43% 1.74%
Spain 231K 6.03% 5.91% 10.69% 2.00%
Poland 214K 6.32% 6.17% 9.48% 0.99%
Australia 199K 7.82% 7.45% 21.55% 1.65%
Italy 199K 5.65% 5.53% 11.99% 1.58%
Netherlands 165K 5.90% 5.57% 10.86% 1.44%
Türkiye 149K 9.82% 9.77% 15.59% 0.85%
Argentina 122K 7.96% 7.47% 11.56% 1.56%
Belgium 107K 5.85% 5.37% 15.05% 1.46%
Mexico 106K 6.04% 6.01% 15.13% 1.60%
Vietnam 105K 11.07% 11.01% 12.36% 2.04%

By platform

Segment Robots.txt known GPTBot ClaudeBot Name any AI crawler Block allBlock all crawlers
Cloudflare 5.36M 14.86% 14.56% 18.58% 2.40%
WordPress 4.94M 5.99% 5.80% 7.21% 0.62%
Shopify 566K 4.62% 4.54% 3.38% 2.99%
Wix 544K 0.83% 0.81% 95.33% 0.08%
Squarespace 299K 2.08% 2.07% 96.65% 0.19%
Joomla 145K 4.20% 4.03% 6.97% 1.00%
Webflow 123K 3.92% 3.63% 6.37% 1.04%
Drupal 103K 6.74% 6.42% 8.15% 1.38%
Magento 43,672 19.58% 19.07% 28.01% 1.58%
Ghost 6,560 7.52% 7.33% 9.51% 0.67%

Robots.txt known counts the segment's domains whose robots.txt policy we could determine; blocked and named percentages divide by that number, and the block-all column divides by the segment's domains with a robots.txt. Authority tiers come from Open Page Rank scores, traffic tiers from the newest Chrome UX Report month, country from where a site's Chrome users are, and platform from our technology detections.

Beyond robots.txt

How to read these numbers

Of the 44.1 million domains in this month's crawl: 30.91 million resolved in DNS, 28.64 million answered our fetch, and 24.91 million gave a definitive robots.txt answer (a readable file or a clean 404, which means everything is allowed). Percentages on this page use that last number unless labeled otherwise.

A robots.txt rule is a published preference, not an enforcement mechanism. We record what sites declare; we cannot see which crawlers comply. Method details on the methodology page.

Look up any website for free

No paywall, no lead-gen, no account. Paste a domain and see what powers it.

No account neededHistory since 2018