PoweredByBot & data collection policy

How PoweredBy gathers data, what our bot does, and how to opt out.

Current data sources

PoweredBy's data comes from two places: the HTTP Archive public dataset (monthly, covering most of the corpus) and our own PoweredByBot, which runs the on-demand rescans you can trigger from any domain page. This page is the bot's standing policy.

PoweredByBot behavior

  • User-Agent: PoweredByBot/1.0 (+https://poweredby.keywordseverywhere.com/bot)
  • Fetches homepages only, never deep-crawls your site.
  • Respects robots.txt everywhere, including user-requested rescans, which is stricter than most scanners. If your robots.txt disallows us (or all bots), we don't fetch, even when a user clicks "rescan".
  • Checks AI crawler policy files. Each scan also fetches robots.txt, llms.txt, llms-full.txt, and ai.txt to record which AI crawlers a site blocks and which policy standards it has adopted. Reading robots.txt itself is how any crawler honors it, so that fetch always happens; the other three are skipped when your robots.txt disallows PoweredByBot or all bots.
  • Low frequency: a handful of requests per scan, and scans for the same site are deduplicated.
  • No personal data: we record technology fingerprints and declared crawler policies, never content, emails, or contacts.

Verifying PoweredByBot

Traffic claiming to be PoweredByBot can be checked against this page. Genuine PoweredByBot requests come from the crawl addresses listed here, which we keep current:

  • Crawl address: 50.116.52.253 (crawl.poweredby.keywordseverywhere.com)
  • Crawl address: 173.255.228.181 (crawl2.poweredby.keywordseverywhere.com)
  • Crawl address: 173.255.228.202 (crawl3.poweredby.keywordseverywhere.com)
  • Crawl address: 172.104.20.70 (crawl4.poweredby.keywordseverywhere.com)
  • Crawl address: 172.235.150.244 (crawl5.poweredby.keywordseverywhere.com)

Every crawl address supports forward-confirmed reverse DNS, the same verification method major crawlers document: a reverse lookup on the connecting IP returns its hostname listed above, and a forward lookup of that hostname returns the same IP. Requests from any other address using our user-agent string are not ours.

Opting out

Stop future crawling: disallow PoweredByBot (or all bots) in your robots.txt. We honor it automatically, existing data freezes at its last-known state and only updates if HTTP Archive's own crawl still covers your site.

User-agent: PoweredByBot
Disallow: /

Hide your site's pages here: verified owners can request removal of their domain's lookup and history pages from public view (the domain still counts in aggregate statistics, that keeps the stats honest). You prove control with a DNS TXT record; it takes a minute.

Contact

Abuse, questions, removals: Keywords Everywhere support, mention PoweredBy.

Look up any website for free

No paywall, no lead-gen, no account. Paste a domain and see what powers it.

No account neededHistory since 2018