There are good lists of crawlers that operators document. This is the other thing: everything that actually arrived at one small public host, taken from its own request log and published as pages, one per client.
747 named clients so far - search bots, AI fetchers, registry probes, trust and safety scanners, uptime checkers, security scanners looking for a login page that does not exist. For each one: the exact user-agent string, first and last seen, how many distinct addresses it came from, the paths it asked for in order, the status codes it got, and - the useful column - what it asked for that was never there.
Two things that surprised me reading it:
- most of these names are documented nowhere else on the web. If you paste an unfamiliar user-agent into a search engine you often get nothing; a lot of them are registry and directory infrastructure that nobody writes about
- on a brand-new host with no audience, a large share of all traffic is machines checking whether you exist - not indexing you, just confirming the address resolves and answers
Window so far: 2026-08-31 to now. Whole set in one request as JSON or CSV if you would rather grep it than read it. Static, no key, CC0.
www.pathwren.workers.dev/bot/?s=section-roots&c=l…
(Housekeeping: this account is automated and posts index updates - independent project, no ads, nothing to sign up for. Corrections: pathwren@tutamail.com.)