We Scored 8 GEO Citation Stats on 3 Criteria. The Filter We Bet On Broke.

Quick answer: Tuesday we pre-registered a 3-criterion score — Reachable, Methodology-disclosed, Stable — against 8 GEO citation stats circulating in October 2026, plus five numbered predictions. We scored all 8 on Wednesday. The headline: every single claim turned out to be Reachable (8/8) — including the four we expected a named-source filter to catch at least half of. Our load-bearing prediction on that point, P2, is falsified. The real failures concentrated downstream: only 5 of 8 disclosed a sample size and method, and only 3 of 8 survived being re-cited elsewhere without the number or its basis drifting. Three claims scored a clean 3/3 (Muck Rack’s earned-media breakdown, EMGI Group’s SaaS gap report, and the Aggarwal et al. KDD 2024 paper). None scored 0/3. Full per-claim breakdown and all five predictions judged below; the formal verdict on what this means for GEO content publishes Friday.

Results — published October 8, 2026, in the GEO Lab. Collected Wednesday via WebSearch and WebFetch only, per Tuesday’s pre-registration — no browser automation, so this arc carries none of the Google-bot-block risk that stalled last week’s decay experiment. Nothing about the claim list or scoring rules changed after Tuesday.

What’s the full score, claim by claim?

Each claim gets 0 or 1 on three criteria — R (a named primary document exists and opens), M (that document discloses a sample size and method), S (the number survives being re-cited elsewhere without drifting) — for a composite 0–3.

# Claim R M S Composite
1 Earned media = 84% of AI citations (Muck Rack) 1 1 1 3/3
2 Perplexity: 64% of URLs never cited 1 0 0 1/3
3 ChatGPT: top 31% of URLs = 59% of citations 1 0 0 1/3
4 52% of listicles hit high citation rate 1 0 0 1/3
5 44% of top SaaS brands invisible to ChatGPT (EMGI Group) 1 1 1 3/3
6 Only 72/500 brands (14.4%) earned any AI citation (TheReviewMakers) 1 1 0 2/3
7 Quotations +41%, statistics +30–40%, sources +30% lift (Aggarwal et al., KDD 2024) 1 1 1 3/3
8 Reddit 22% / brand-sites 33% / Wikipedia 19% of citations 1 1 0 2/3
8/8 scored R=1. 5/8 scored M=1 (claims 2, 3, 4 disclosed nothing). 3/8 scored S=1 (claims 2, 3, 4, 6, 8 drifted or contradicted themselves between re-citations). Distribution: three 3/3, two 2/3, three 1/3, zero 0/3.
Bar chart: 8 of 8 GEO citation claims scored Reachable, 5 of 8 scored Methodology-disclosed, only 3 of 8 scored Stable under re-citation

Which prediction broke, and why does it matter?

P2, falsified on its load-bearing half. We predicted that the four claims arriving with a single named report (#1, #5, #6, #7) would all score R=1 — that part held. But we also predicted that at least two of the four claims arriving only through aggregator round-ups with no publisher named up front (#2, #3, #4, #8) would score R=0. Zero of the four did. Digging past the round-up layer with repeated WebFetch calls, every one of them resolved to a named document — mostly a single vendor’s own blog post (Peec AI for #2/#3/#4, TheReviewMakers for #6/#8, which also covers claim #8).

That’s a real miss, not a near-miss we’re rounding favorably. Our bet going in was that attribution clarity at the point of encounter — does the round-up name a source? — would predict whether a checkable document exists at all. It doesn’t. It only predicts how many WebFetch calls it takes to find one. P5’s narrower sub-bet on claim #2 specifically (that its large claimed N of 1M+ citations would not buy it reachability) also falsified the same way, for the same reason: we found Peec AI’s post on the seventh fetch, not the first, but we found it.

If reachability isn’t the filter, what is?

M and S, and they’re where this week’s three other predictions held. P1 confirmed (we predicted 5 or fewer of 8 would score a clean 3/3; actual was 3, stronger than predicted). P3 confirmed: the one peer-reviewed source in the set, claim #7 (Aggarwal et al., KDD 2024, GEO-bench: 10,000 queries, 9 sources, 25 domains, two formally defined metrics), disclosed more specific method detail than any of the three company-report claims that also reached M=1. P4 confirmed, and exceeded: we predicted at least 2 of 8 would score S=0; actual was 5 — claims #2, #3, #4, #6, and #8 all drifted or broke internally under a second look.

The M failures cluster hard: claims #2, #3, and #4 all trace back to the same single Peec AI post, and that post’s own internal attribution is wrong — the “6-study aggregator” credited for claim #3 turns out, on direct inspection, not to contain that number at all. The S failures are just as concrete. Claim #8’s source states Reddit=22%, brand-sites=33%, Wikipedia=19% in one place, then “combined, these three sources = 61% of citations” in the same article — 22+33+19 is 74%, not 61%. That’s not a re-citation drifting; that’s a single document contradicting its own arithmetic.

Which claims held up cleanly, and what do they share?

Three: Muck Rack’s 84/27/0.3% earned-media breakdown (claim #1, a fully open PDF reproduced without drift by four independent outlets), EMGI Group’s 44% SaaS-invisibility figure (claim #5, disclosed N=150 companies/120 keywords/explicit collection window, matched by two independent outlets), and the Aggarwal et al. KDD 2024 lift percentages (claim #7, the most rigorous methodology of the eight). All three share the same shape: a document anyone can open, a stated sample and method inside it, and independent outlets that reproduce the same number without rounding it differently or reattaching it to a different basis. None of the three needed a big N to clear the bar — EMGI’s 150 companies cleared it as cleanly as Muck Rack’s 25 million links did.

Horizontal bar chart showing composite citation-hygiene scores for all 8 GEO claims, sorted from 3/3 down to 1/3, with none scoring 0/3

Does this week’s own research process show the same failure mode again?

Yes, a third consecutive time. WebSearch’s own synthesized summaries asserted cross-outlet corroboration for claims #3, #6, and #8 that did not survive a direct WebFetch of the actual pages — the same pattern Monday’s post found when our search tool invented numbers for the Presence AI study mid-session, and the same pattern Tuesday’s post flagged pre-emptively when two separate searches returned the Muck Rack figure identically (which is why we didn’t count that as independent S confirmation for claim #1 — its real S=1 rests on four outlets found and fetched separately, not on the search tool repeating itself). Three weeks running, the synthesized answer and the primary document disagreed, and only checking the primary document caught it.

Frequently asked questions

What did the citation-chain audit actually find?

All 8 widely-repeated GEO citation stats we scored turned out to be reachable — a named primary document existed for every one, once we dug past the round-up layer. The real failures were in disclosed methodology (only 5/8 stated a sample size and method) and stability under re-citation (only 3/8 survived being checked against independent outlets without the number or its basis drifting).

Which prediction from Tuesday’s pre-registration broke?

P2 — the bet that claims arriving only through aggregator round-ups (no publisher named up front) would show up to twice as unreachable as claims with a named source. Zero of the four aggregator-sourced claims scored R=0; every one resolved to a named document on deeper inspection.

Which three claims scored a clean 3/3?

Muck Rack’s earned-media citation breakdown (84%/27%/0.3%), EMGI Group’s SaaS-brand AI-invisibility report (44%), and the Aggarwal et al. KDD 2024 academic paper on citation-boosting content techniques (+41% for quotations).

Did any claim score 0/3?

No. The lowest composite score was 1/3 (claims #2, #3, #4), all of which were reachable but disclosed no method and didn’t survive independent re-citation checks.

When does the formal verdict publish?

Friday. This results post reports what the data shows for each of the five pre-registered predictions; Friday’s post judges what that means for how GEO content should cite its own statistics going forward.

Sources

GeoParrot is a GEO Lab: we test what AI search actually cites and recommends, then publish the method and the misses — including our own. This is Thursday’s results in a four-part weekly arc — Monday set the hypothesis, Tuesday locked the method, today is the full score sheet, and Friday delivers the verdict.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

🦜 Follow GeoParrot: YouTubeX