Content-Signal adoption
Content signals are machine-readable usage preferences carried inside
robots.txt: search, ai-train, and ai-input, each set to
yes or no. The format was introduced by Cloudflare in 2025 and can be injected automatically
for sites using its managed robots.txt.
ai-train=no
0.02%
say no to AI input
3,676 domains declare ai-input=no, no AI answers from their content
<0.01%
say no to search
127 domains declare search=no, rare by design
Data for Aug 2026, from our monthly crawl. Every percentage divides by the base the first tile names.
Notable adopters
- cloudflare.com 9.8
- cookiebot.com 9.2
- pewresearch.org 9.1
- cookielaw.org 9.1
- usercentrics.eu 9.1
- bizjournals.com 9.0
- mayoclinic.org 9.0
- mdpi.com 9.0
- theatlantic.com 9.0
- hugedomains.com 9.0
- uk.com 9.0
- freepik.com 9.0
- patreon.com 9.0
- utexas.edu 8.9
- webflow.com 8.9
- zapier.com 8.9
- jsfiddle.net 8.9
- supabase.co 8.9
- nami.org 8.9
- podbean.com 8.9
- maryland.gov 8.8
- it.com 8.8
- hud.gov 8.8
- gitbook.io 8.8
Ranked by Open Page Rank. The per-domain page shows each site's full AI crawler policy.
Where this standard actually stands
Content signals are new and their reach mirrors Cloudflare's: a large share of declarations arrive through Cloudflare's managed robots.txt rather than hand-edited files, so adoption measures platform rollout as much as individual publisher choice. Whether AI operators honor the signals is a separate question that a robots.txt file cannot answer.
Counting method and definitions are on the methodology page.