The verdict: no verdict. Tuesday we pre-registered five predictions on whether an AI citation you win today survives to next week, with a clean bet — “rented, not owned, but the rent has a floor.” We never got to test it. The first scheduled snapshot, Day 0, hung on a blank page: Google’s search surface returned readyState=loading forever, an empty document, no error. We held off re-trying to avoid flagging the automation profile further. Today, four days later, we tried once more with a single query — same result, 0 characters of content, confirming this isn’t a one-off glitch. So the five predictions (P1–P5) stay exactly as locked: untested, not confirmed, not rejected. We’re publishing that null result instead of quietly dropping the week, because the failure itself is a finding: the same anti-automation wall that choked rank trackers by ~80% in mid-September also blocks a legitimate, logged-out research script from querying Google’s AI Mode — and it fails silently, as an empty page rather than an error, which is exactly the failure mode that could make a careless citation-monitoring tool report “you lost your citation” when the truth is “we got blocked.”
Friday verdict — published October 2, 2026, closing Week 15 of the GEO Lab. The arc: last week’s results found that re-running an identical query on Google AI Mode kept only 53% of cited domains within the same session — a noise floor. Tuesday we pre-registered a daily-snapshot test to see if citations erode further over days, not minutes. Wednesday’s collection attempt is the subject of this post.
What actually happened when we tried to collect Day 0?
Wednesday morning, the collector opened the same automated browser profile we’ve used for every experiment this series — the one that produced last week’s 0.53 noise-floor number and the cross-engine citation study before it — and pointed it at a Google AI Mode query (google.com/search?udm=50). The page never finished loading. document.readyState stayed at loading indefinitely, and the rendered document was empty — no citation panel, no answer text, nothing to parse. We ruled out a network problem first: the same browser rendered example.com and geoparrot.com normally in the same session, and a plain shell curl to google.com returned a clean HTTP 200. The page loaded for everyone except that one Chrome profile hitting that one surface. That’s the signature of a soft bot block — Google recognizing the remote-debugging profile as automation and quietly withholding real content rather than returning a 403 or CAPTCHA you could at least detect and handle.
We stopped there on purpose. Retrying aggressively into a block is how you turn a soft flag into a permanent one, so Day 0 stayed unscored and we logged the failure instead of forcing it. Today, October 2 — four days and zero retries later — we ran one more single-query check as a diagnostic for this post, not as data collection. Same result: the AI Mode panel returned 0 characters of answer text and 0 cited domains for “best vpn for streaming while traveling,” one of the eight frozen queries from Tuesday’s design. Whatever flagged the profile on Wednesday hasn’t lifted. This wasn’t a transient outage; it’s a standing block on this collection method.
So what’s the verdict on the five predictions?
There isn’t one, and we’re not going to manufacture one. Pre-registration only means something if a missed collection gets reported as missed, not quietly reframed as a result. Here’s exactly where each prediction stands:
| Prediction | Status | Why |
|---|---|---|
| P1 (load-bearing) — temporal overlap falls below the 0.53 noise floor but above 0.20 | ⬜ Untested | No Day 0 baseline exists, so there’s nothing to compare any later day against. |
| P2 — an invariant core domain exists for most queries | ⬜ Untested | Requires multiple daily snapshots; we have zero. |
| P3 — the core is authority-weighted | ⬜ Untested | Same dependency as P2. |
| P4 — decay is monotonic across lag | ⬜ Untested | Needs at least three data points over time; collection never reached one. |
| P5 — Reddit/YouTube persist as invariant anchors | ⬜ Untested | Same dependency as P2–P4. |
We considered two shortcuts and rejected both. First: switch engines mid-week to ChatGPT or Perplexity, which have been more automatable in earlier rounds, and quietly collect a decay curve there instead. We said no — Tuesday’s design locked Google AI Mode specifically for precision, and swapping the instrument after the pre-registration is the exact move pre-registration exists to prevent. Second: treat the single Day-0 failure plus today’s retry as “the citation panel evaporated to zero,” which would make P1 look spectacularly confirmed. That’s a measurement artifact wearing a result’s clothes, and we’re not laundering it into a headline.
Doesn’t the block itself tell us something?
Yes — and it’s arguably the more useful finding for most readers than the decay number would have been. Two weeks ago we covered Google’s anti-scraping crackdown that cut third-party rank-tracker data by roughly 80% starting around September 13, and we made the point that most “AI visibility” dashboards scrape the same Google surface to tell you whether you appear in AI Overviews and AI Mode. This week we became the case study: a legitimate, logged-out, slow, pre-registered research script — not an aggressive scraper hammering the API — still got silently blocked at the first query of a brand-new experiment. The enforcement isn’t just throttling bulk rank-tracking traffic anymore; it’s catching individual automated browser sessions that query Google’s AI surface at any volume.
The failure mode is the part worth sitting with. The block didn’t return an error, a CAPTCHA, or a rate-limit message — any of which a careful tool could detect and flag. It returned a page that looks loaded (valid HTML, a real DOM) but carries zero citation data. A citation-monitoring tool that isn’t specifically checking for empty-content-on-a-200-response would log that as “0 domains cited today” and hand you a chart showing you lost every citation you had — when what actually happened is the tool got quietly defended against, not you. That’s the inferred-metric problem we flagged in “counted vs. inferred” made concrete: a number built on scraping a surface you don’t control can go to zero for reasons that have nothing to do with your actual visibility, and nothing in the output tells you which happened.
What should you do if you rely on an AI-citation tracker?
A practical checklist, built directly from hitting this wall ourselves:
- If a vendor’s “AI visibility” number scrapes Google AI Mode or AI Overviews: ask them directly how they distinguish “zero citations” from “collection blocked.” If they can’t answer, assume some fraction of every zero or sudden drop in their history is a silent block, not a real loss.
- If you’re building your own monitor via browser automation: expect this exact failure mode — a 200-status page with an empty document, not a clean error — and add a content-sanity check (character count, DOM-element presence) before you trust any “0 results” reading. Treat remote-debugging CDP profiles as a liability; they appear to be a specific flag target.
- Lean on counted, first-party signals for anything decision-critical. Search Console’s Generative AI report comes from Google’s own logs, not scraping, so this class of failure can’t touch it — it’s the one number in this whole saga that didn’t need a bot-detection workaround to exist.
- Don’t let a single blocked engine stop your whole measurement program. This block is specific to automated querying of Google’s search surface. Our cross-engine testing already established the four major engines barely overlap in what they cite — a Google-side block tells you nothing about ChatGPT or Perplexity coverage.
- Pre-register before you collect, publicly, every time. The entire reason this post exists instead of a quietly-abandoned experiment is that Tuesday’s predictions were already live before Wednesday’s collection failed. If we hadn’t pinned them, this week would have just vanished. That discipline cost us nothing and caught the thing worth reporting.
What are we testing next?
The decay question is still open, and we’re not dropping it — we’re changing the instrument. Next round we’ll rebuild the same eight-query, Day-0-vs-later-day design around a collection method that degrades visibly instead of silently: wider spacing between snapshots, a content-sanity check that halts and flags instead of recording a zero, and a fallback to manual spot-checks if automation gets flagged again. If Google’s defenses make programmatic daily monitoring of its AI surface structurally unreliable for independent researchers — not just us — that’s worth establishing on its own, separate from whatever the decay number turns out to be. Follow the GEO Lab for the rebuilt attempt.
Frequently asked questions
So do AI citations decay over time or not?
We don’t know yet. The pre-registered experiment designed to answer this never collected its Day 0 baseline — Google’s bot defenses blocked the automated query before a single snapshot was scored. All five predictions from Tuesday’s design remain untested, not confirmed and not rejected. We’re rebuilding the collection method for another attempt rather than guessing.
What exactly blocked the data collection?
Google’s search results page, including AI Mode, hung at readyState=loading with an empty document for our automated browser profile, while the same browser rendered other sites normally and a plain network request to google.com returned HTTP 200. That pattern points to a soft bot-detection block on the automation profile specifically — content withheld, not a network failure — and it was still present on a retry four days later.
Why not just use the failure as proof citations collapsed to zero?
Because that would conflate a measurement artifact with a real result. A blocked collector returning zero domains looks identical, on the surface, to a citation panel that genuinely emptied out — but one is about our tooling and the other would be about the engine’s actual behavior. Reporting the first as the second is exactly the “inferred metric” failure mode we’ve warned about, so we’re keeping them separate.
Does this affect tracking on ChatGPT or Perplexity too?
Not based on what we observed this week. The block we hit is specific to automated querying of Google’s search surface (including AI Mode), the same surface implicated in September’s rank-tracker scraping crackdown. ChatGPT and Perplexity query through separate pipelines; our earlier collection rounds on those engines didn’t show this failure mode.
What should I check if my AI-citation tracker shows a sudden drop to zero?
Before treating it as a real loss, ask whether the tool can distinguish “no citations found” from “collection blocked” — a silently blocked scraper and a genuinely empty result set look the same in most dashboards. Cross-check against a signal that doesn’t depend on scraping, like Search Console’s Generative AI report, before reacting to the number.

Leave a Reply