Websites blocking CCBot

The crawler of the Common Crawl nonprofit, which publishes a free archive of web data. That archive is one of the most widely used sources of LLM training data, so blocking CCBot is a common proxy for opting out of AI training generally.

Operated by Common Crawl · Training crawler · operator documentation

10.01% block CCBot 2.49 million of 24.91 million domains with a known robots.txt policy 6.63% block it by name 1.65 million domains, the rest inherit a blanket * block 2.50% restrict it by name 624K domains name it and close some of its paths 9.53% name it in robots.txt 2.37 million domains give it rules of their own

Data for Aug 2026. All four percentages divide by the same 24.91 million domains. Blocked counts both explicit User-agent: CCBot blocks and blanket * blocks the crawler inherits.

Historical mentions, HTTP Archive

Share of sites in each month's HTTP Archive crawl whose robots.txt named CCBot. This series counts mentions only (the archive keeps no rule text), comes from a different population than our own crawl above, and the two are never merged.

High-authority sites blocking CCBot by name

Domain OPROpen Page Rank Blocking since
linkedin.com 9.82/10 Aug 2026
vimeo.com 9.54/10 Aug 2026
nytimes.com 9.28/10 Aug 2026
tumblr.com 9.27/10 Aug 2026
washingtonpost.com 9.26/10 Aug 2026
who.int 9.26/10 Aug 2026
bbc.co.uk 9.18/10 Aug 2026
bbc.com 9.15/10 Aug 2026
cookielaw.org 9.10/10 Aug 2026
usatoday.com 9.08/10 Aug 2026
sagepub.com 9.07/10 Aug 2026
bandcamp.com 9.06/10 Aug 2026
tripadvisor.com 9.05/10 Aug 2026
lnkd.in 9.05/10 Aug 2026
congress.gov 9.05/10 Aug 2026
bizjournals.com 9.04/10 Aug 2026
msn.com 9.04/10 Aug 2026
mayoclinic.org 9.04/10 Aug 2026
pexels.com 9.03/10 Aug 2026
theatlantic.com 9.01/10 Aug 2026
hugedomains.com 9.00/10 Aug 2026
uk.com 9.00/10 Aug 2026
patreon.com 8.99/10 Aug 2026
healthline.com 8.98/10 Aug 2026
media-amazon.com 8.95/10 Aug 2026
ebay.com 8.95/10 Aug 2026
science.org 8.95/10 Aug 2026
lefigaro.fr 8.95/10 Aug 2026
usgs.gov 8.95/10 Aug 2026
utexas.edu 8.92/10 Aug 2026

Sites whose robots.txt names CCBot (or a legacy alias) with a full block, ranked by Open Page Rank. Sites blocking every crawler with a blanket * rule are counted above but not listed here. Since dates start at our first observation of the rule.

Restricting some paths

Look up any website for free

No paywall, no lead-gen, no account. Paste a domain and see what powers it.

No account neededHistory since 2018