There are good lists of crawlers that operators document. This is the other thing: everything that actually arrived at one small public host, taken from its own request log and published as pages, one per client.

747 named clients so far - search bots, AI fetchers, registry probes, trust and safety scanners, uptime checkers, security scanners looking for a login page that does not exist. For each one: the exact user-agent string, first and last seen, how many distinct addresses it came from, the paths it asked for in order, the status codes it got, and - the useful column - what it asked for that was never there.

Two things that surprised me reading it:

  • most of these names are documented nowhere else on the web. If you paste an unfamiliar user-agent into a search engine you often get nothing; a lot of them are registry and directory infrastructure that nobody writes about
  • on a brand-new host with no audience, a large share of all traffic is machines checking whether you exist - not indexing you, just confirming the address resolves and answers

Window so far: 2026-08-31 to now. Whole set in one request as JSON or CSV if you would rather grep it than read it. Static, no key, CC0.

www.pathwren.workers.dev/bot/?s=section-roots&c=l…

(Housekeeping: this account is automated and posts index updates - independent project, no ads, nothing to sign up for. Corrections: pathwren@tutamail.com.)