Websites blocking Webzio-Extended

Webz.io's opt-out token for AI training uses of its crawled data. Webz.io sells web datasets to AI companies.

Operated by Webz.io · Training crawler

4.09% block Webzio-Extended 1.02 million of 24.91 million domains with a known robots.txt policy 0.69% block it by name 173K domains, the rest inherit a blanket * block 2.31% restrict it by name 576K domains name it and close some of its paths 3.03% name it in robots.txt 755K domains give it rules of their own

Data for Aug 2026. All four percentages divide by the same 24.91 million domains. Blocked counts both explicit User-agent: Webzio-Extended blocks and blanket * blocks the crawler inherits.

Historical mentions, HTTP Archive

Share of sites in each month's HTTP Archive crawl whose robots.txt named Webzio-Extended (or a legacy alias of it). This series counts mentions only (the archive keeps no rule text), comes from a different population than our own crawl above, and the two are never merged.

High-authority sites blocking Webzio-Extended by name

Domain OPROpen Page Rank Blocking since
usatoday.com 9.08/10 Aug 2026
congress.gov 9.05/10 Aug 2026
msn.com 9.04/10 Aug 2026
investopedia.com 9.01/10 Aug 2026
healthline.com 8.98/10 Aug 2026
lefigaro.fr 8.95/10 Aug 2026
newyorker.com 8.90/10 Aug 2026
mashable.com 8.87/10 Aug 2026
smh.com.au 8.86/10 Aug 2026
arstechnica.com 8.84/10 Aug 2026
medicalnewstoday.com 8.83/10 Aug 2026
gizmodo.com 8.82/10 Aug 2026
wikihow.com 8.81/10 Aug 2026
pcmag.com 8.80/10 Aug 2026
slate.com 8.80/10 Aug 2026
amazon.it 8.78/10 Aug 2026
adweek.com 8.77/10 Aug 2026
fliphtml5.com 8.75/10 Aug 2026
verywellmind.com 8.72/10 Aug 2026
metro.co.uk 8.72/10 Aug 2026
tagesspiegel.de 8.71/10 Aug 2026
people.com 8.70/10 Aug 2026
sfgate.com 8.70/10 Aug 2026
mainichi.jp 8.69/10 Aug 2026
nikkeibp.co.jp 8.68/10 Aug 2026
nos.nl 8.67/10 Aug 2026
launchpad.net 8.63/10 Aug 2026
computerworld.com 8.63/10 Aug 2026
vermont.gov 8.63/10 Aug 2026
fedoraproject.org 8.62/10 Aug 2026

Sites whose robots.txt names Webzio-Extended (or a legacy alias) with a full block, ranked by Open Page Rank. Sites blocking every crawler with a blanket * rule are counted above but not listed here. Since dates start at our first observation of the rule.

Restricting some paths

Also matched via the legacy token omgilibot, which some sites still use for this crawler.

Look up any website for free

No paywall, no lead-gen, no account. Paste a domain and see what powers it.

No account neededHistory since 2018