This Week, Two Sites Published Raw AI-Crawler Logs. Combined Fetches of llms.txt: Zero.

Quick answer: This week two site operators independently published raw AI-crawler logs, verified against the IP ranges vendors publish themselves (not user-agent strings, which can be spoofed). Prosopo (Oct 5, 2026) logged 4 days of traffic to a site that’s hosted an llms.txt since May 2026: robots.txt fetched 120 times by verified AI crawlers, sitemap files 43 times, llms.txt and llms-full.txt — zero. A separate dev.to writeup (Oct 6–7, 2026) analyzing 2,849 requests to a new tools-and-tutorials site found the same robots.txt/sitemap pattern and didn’t log a single llms.txt fetch either (the site doesn’t appear to have one, so this is corroborating by absence, not a second zero-count). Combined verified-crawler fetch count for llms.txt across both logs: 0. That’s a third independent method — raw IP-verified server logs, this time from people with no stake in the GEO industry’s answer — landing on the same conclusion as our own 137,000-site study and Google’s own June confirmation.

Trend watch — published October 7, 2026. We read every new llms.txt data point the way the GEO Lab reads any claim: what was actually measured, how, and whether it replicates.

What exactly did these two logs measure, and how solid is the method?

Prosopo — a bot-detection company, so reading crawler logs is literally their business — published four days of IP-verified traffic to a site that’s maintained an llms.txt (19 KB) and llms-full.txt (370 KB) since May 2026. Their method: match each request’s source IP against the crawler IP ranges OpenAI, Anthropic, Perplexity, and DuckDuckGo publish themselves, rather than trusting the User-Agent header, which any scraper can fake. Under that stricter standard, verified AI crawlers hit robots.txt 120 times (Anthropic 47, OpenAI 31, DuckAssist 27, Perplexity 15) and sitemap files 43 times (Anthropic 40, OpenAI 3). llms.txt and llms-full.txt: zero verified hits. There were 40 unverified requests to the llms files, which Prosopo attributes to vulnerability scanners probing for exposed config files, not AI agents reading documentation.

The dev.to piece, published a day later, took a different site and a different angle: 2,849 total requests to a brand-new site (30 HTML tools, a 31-post blog, 4 SEO tutorials) on October 6, broken down by bot identity. Of 321 bot requests, 183 (57%) came from AI-labeled crawlers — ClaudeBot (49), GPTBot (20), plus smaller AI-research bots. robots.txt got 35 fetches, the homepage 17, sitemap.xml 16. This site doesn’t appear to run an llms.txt, so it can’t add a second zero-count to the tally — what it adds is a look at what AI crawlers actually read once they show up, which turned out to be concentrated on the site’s tutorial pages (50 combined fetches) far more than its tool pages, which the author describes as mostly “shell and metadata” to a bot.

Signal Prosopo (4-day log, IP-verified) dev.to site (Oct 6, 2,849 requests)
robots.txt fetches 120 (verified AI crawlers) 35
sitemap fetches 43 (verified AI crawlers) 16
llms.txt / llms-full.txt fetches 0 (verified) — 40 unverified, likely scanners not applicable (no llms.txt)
Content AI crawlers favored not broken out tutorial/explainer pages over tool pages

Why does one more zero-count matter when we already had two?

Because the first two data points came with an obvious objection: our 137,000-site study and Google’s own statement both came from parties with something to prove about llms.txt — us by running a GEO Lab that’s skeptical of unverified tactics, Google by having a competing interest in not legitimizing a standard it doesn’t control. Prosopo has no dog in the GEO fight; they sell bot detection and published their logs as a methodology flex about IP verification, with the llms.txt result almost a side note. That’s a materially different incentive structure landing on the same zero. By the scoring method we pre-registered this week for auditing GEO stats — reach, methodology disclosed, re-citation stability — this evidence scores well on the methodology axis specifically because it names its verification method (IP ranges, not user-agent) and shows its raw counts instead of a rounded claim.

If crawlers skip llms.txt, what do they reliably read?

Both logs agree on this part: robots.txt and sitemap.xml get fetched, every time, by every major AI crawler brand. Those are the only discovery files with confirmed vendor support — not a documentation standard nobody committed to reading. The dev.to breakdown adds a second, more actionable signal: on a brand-new site, AI crawlers spent disproportionate attention on tutorial and explainer content relative to tool pages, because tutorial pages are — in the author’s words — “text-heavy, semantically dense, and link-rich,” while a tool page is mostly interactive shell with little parseable text. That’s one new site, so treat it as a hypothesis, not a rule. But it points in the same direction as findings we’ve already reported: first-party substance outperforms markup, and markup-only moves (schema, llms.txt, meta tricks) keep underperforming content AI can actually parse and quote.

A checklist if you’re deciding what to do with llms.txt right now

  • Don’t expect llms.txt to drive AI-citation traffic. Three independent methods — our 137K-site study, Google’s own confirmation, and now two IP-verified crawler logs — agree no major AI crawler reliably fetches it.
  • Keep (or skip) it for the reason that’s actually true: optional human/agent-SDK readability, not a citation lever. If you already have one, it’s not hurting you; it’s just not the growth channel some vendors pitch.
  • Make sure robots.txt isn’t accidentally blocking the crawlers you want. It’s the one file every verified AI crawler actually reads — see our top-5,000 robots.txt census for how often sites get this wrong by mistake.
  • Spend the effort on text-dense, link-rich explainer content instead of thin tool/product shells. That’s where both logs saw crawlers actually linger.
  • Weight vendor IP-range verification over User-Agent claims in any crawler-log study you read — including this one. User-Agent strings are trivially spoofed; IP-range matching against vendor-published ranges is the harder, more credible standard.

The honest caveats

Two more data points is not a census. Prosopo’s log covers four days on one site; the dev.to log covers one day on one brand-new site with zero indexed pages at the time. Site operators who bother to publish raw crawler logs and write about llms.txt are self-selected toward people already skeptical of unverified GEO claims, which could bias the sample toward the “it doesn’t work” conclusion we also hold. Neither log can rule out that some AI crawler reads llms.txt on other sites, under other conditions, or that vendors start supporting it later — absence of evidence across roughly 7 combined days of two sites isn’t proof of permanent absence everywhere. What we can say is that every independent measurement method tried so far — large-scale aggregate study, a vendor’s own statement, and now two separate IP-verified server logs — has landed on the same zero, which is the kind of convergence across different methods that should move a prior, even without a true census.

Frequently asked questions

Did two sites really get zero llms.txt fetches this week?

One did directly: Prosopo logged zero verified AI-crawler fetches of llms.txt/llms-full.txt across 4 days, despite hosting the file since May 2026. The second site (analyzed on dev.to) doesn’t run an llms.txt, so it couldn’t log a fetch either way — it corroborates the robots.txt/sitemap pattern instead.

How is this different from your June study or Google’s statement?

Different method, same answer. Our 137,000-site study measured adoption and correlation at scale; Google made a direct policy statement; these two logs are raw, IP-verified server traffic from independent, non-GEO-industry site operators. Three different methods converging is stronger evidence than any one of them alone.

What’s the difference between IP-range verification and User-Agent checking?

User-Agent is a header any request can set to anything, including “GPTBot” when it isn’t. IP-range verification checks the request’s actual source IP against ranges OpenAI, Anthropic, Perplexity, and others publish as belonging to their real crawlers — much harder to fake.

Should I remove my llms.txt file?

No strong reason to. The data says it won’t drive AI-citation traffic, not that it causes harm. Keep it if you want agent/developer-readability; don’t keep investing time in it expecting a citation payoff.

What should I actually spend that effort on instead?

Based on the same logs: make sure robots.txt isn’t blocking crawlers you want, and build text-dense, link-rich explainer or tutorial content rather than thin tool/product pages — that’s where crawlers in both studies spent the most time reading.

GeoParrot is a GEO Lab: we run experiments to test what AI search engines actually reward, and we grade the industry’s claims — including our own — against what was actually measured. See the running scoreboard at our 2026 GEO benchmark. Related: our 137,000-site llms.txt study, Google’s confirmation it doesn’t use llms.txt, our robots.txt census, and our method for auditing GEO stats.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

🦜 Follow GeoParrot: YouTubeX