Websites blocking cohere-training-data-crawler
Cohere's crawler for collecting web data used to train its enterprise language models.
Operated by Cohere · Training crawler
* block
2.38%
restrict it by name
592K domains name it and close some of its paths
3.27%
name it in robots.txt
815K domains give it rules of their own
Data for Aug 2026. All four percentages
divide by the same 24.91 million domains. Blocked counts both
explicit User-agent: cohere-training-data-crawler blocks and blanket
* blocks the crawler inherits.
Historical mentions, HTTP Archive
High-authority sites blocking cohere-training-data-crawler by name
| Domain | OPROpen Page Rank | Blocking since |
|---|---|---|
| usatoday.com | 9.08/10 | Aug 2026 |
| congress.gov | 9.05/10 | Aug 2026 |
| msn.com | 9.04/10 | Aug 2026 |
| investopedia.com | 9.01/10 | Aug 2026 |
| newyorker.com | 8.90/10 | Aug 2026 |
| mashable.com | 8.87/10 | Aug 2026 |
| francetvinfo.fr | 8.87/10 | Aug 2026 |
| thelancet.com | 8.85/10 | Aug 2026 |
| note.com | 8.84/10 | Aug 2026 |
| arstechnica.com | 8.84/10 | Aug 2026 |
| pcmag.com | 8.80/10 | Aug 2026 |
| slate.com | 8.80/10 | Aug 2026 |
| amazon.it | 8.78/10 | Aug 2026 |
| adweek.com | 8.77/10 | Aug 2026 |
| ardmediathek.de | 8.77/10 | Aug 2026 |
| fliphtml5.com | 8.75/10 | Aug 2026 |
| tagesschau.de | 8.73/10 | Aug 2026 |
| rfi.fr | 8.73/10 | Aug 2026 |
| verywellmind.com | 8.72/10 | Aug 2026 |
| tagesspiegel.de | 8.71/10 | Aug 2026 |
| people.com | 8.70/10 | Aug 2026 |
| ndr.de | 8.70/10 | Aug 2026 |
| mainichi.jp | 8.69/10 | Aug 2026 |
| france24.com | 8.68/10 | Aug 2026 |
| cell.com | 8.68/10 | Aug 2026 |
| nikkeibp.co.jp | 8.68/10 | Aug 2026 |
| wdr.de | 8.67/10 | Aug 2026 |
| launchpad.net | 8.63/10 | Aug 2026 |
| computerworld.com | 8.63/10 | Aug 2026 |
| vermont.gov | 8.63/10 | Aug 2026 |
Sites whose robots.txt names cohere-training-data-crawler (or a legacy
alias) with a full block, ranked by Open Page Rank. Sites blocking every crawler with a blanket
* rule are counted above but not listed here. Since dates start at our first
observation of the rule.
Restricting some paths
- statista.com
- squarespace.com
- frontiersin.org
- theverge.com
- vox.com
- newsweek.com
- buzzfeed.com
- youradchoices.ca
- ccm19.de
- ign.com
- datareportal.com
- eater.com
Also matched via
the legacy token cohere-ai,
which some sites still use for this crawler.