The verdict, in one line. ✅ The rank-decoupled thesis passes. Across 8 buyer-intent queries, the share of an AI engine’s cited domains that also rank in Google’s top-10 came to 34.4% blended — under our pre-registered ≤50% ceiling — with a 40.3-point spread between Google AI Mode (53.6%) and ChatGPT (13.3%), well over our ≥25-point floor. Both load-bearing criteria clear. We do not retract. The one correction we owe Monday’s fact-check: the viral “under 20%” figure is a ChatGPT number, not a universal one — on Google AI Mode, ranking still buys you roughly half your citations.
What exactly were we ruling on?
On Monday we fact-checked a widely-shared claim: that the overlap between a top Google ranking and the sources AI engines cite collapsed from roughly 70% to under 20%. We said the collapse is real but neither zero nor uniform across engines. On Tuesday we refused to leave it at other people’s aggregates and pre-registered a blind test — writing down the predicted overlap bands before collecting anything, so this write-up could not move the goalposts.
The frozen protocol: 8 high-commercial-intent “best [category]” queries — CRM, project management, email marketing, WordPress hosting, VPN, password manager, freelance accounting, AI writing — the exact SERPs where Google’s top-10 is dominated by affiliate listicles (PCMag, Forbes Advisor, TechRadar). For each, we pulled Google’s top-10 organic domains and the distinct domains cited by Google AI Mode, Perplexity, and ChatGPT (web search on). Overlap = the share of an engine’s cited domains that also appear in that query’s Google top-10, averaged across queries, excluding Google’s own properties and each engine’s self-citation. The bands we locked:
- Google AI Mode 55–75% — generated on top of Google’s own index, so it should stay closest to the rankings.
- Perplexity 25–45% — independent retrieval, leans on reviews and community.
- ChatGPT 15–35% — the loosest coupling to classic rank.
- Blended ≤50% and engine spread (AI Mode − ChatGPT) ≥25 points — the two load-bearing criteria.
The rule we committed to: the rank-decoupled thesis passes only if blended ≤50% AND spread ≥25 points. We retract if blended tops 55% and the engines move together (spread <15). Here is what the blind run returned.
Did the rank-decoupled thesis pass its own pre-registered test? ✅ Yes — on both criteria

Measured overlap, per engine: Google AI Mode 53.6% (8/8 queries), Perplexity 36.4% (7/8), ChatGPT 13.3% (3/8 — more on that limit below). That produces a blended overlap of 34.4% and a spread of 40.3 points. Against the frozen rule:
- Criterion 1 — blended ≤50%: 34.4% ✅ (comfortably clear)
- Criterion 2 — spread ≥25 points: 40.3 ✅ (well clear)
- Retract condition (blended >55% and spread <15): not remotely triggered ❌

The structure Monday described held on our own data: AI answers are not a re-skin of Google’s top-10. On a buyer query, the average AI engine sourced roughly two-thirds of its citations from domains that do not rank on Google’s first page — and the engines disagreed with each other by 40 points on how much rank still matters. The “rank-mirror” model loses.
So is Monday’s “70% → under 20%” claim confirmed? ✅ Conditionally — with one correction
Directionally, yes. Overlap sits far below the old ~70% regime, it is decidedly non-uniform (a 40-point spread), and it is not zero — there is a positive tail on every engine, exactly what Monday’s r ≈ 0.18 correlation predicted. Rank still buys the ticket; it no longer guarantees the seat.
But the precise viral number needs a correction we can now make from first-hand data: “under 20%” describes ChatGPT, not the market. ChatGPT’s overlap (13.3%) is the only engine that lands anywhere near that headline. The blended reality is 34.4%, and on Google AI Mode it is 53.6% — meaning a #1 ranking still lands you in AI Mode’s sources more than half the time. Quoting “under 20%” as if it were the state of AI search cherry-picks the single most decoupled engine and hides the fact that on Google’s own generative surface, classic SEO is still doing most of the work. This is the same lesson as our no-universal-GEO-strategy finding: a single blended number lies; the per-engine spread is the story.
Where did our own predictions miss — and why does the direction matter?
Honesty check on ourselves. Only one of three point estimates landed inside its band: Perplexity at 36.4% (predicted 25–45%). Both others fell just outside — Google AI Mode at 53.6% (predicted 55–75%) and ChatGPT at 13.3% (predicted 15–35%). Crucially, both misses are on the same side: low. Neither engine overlapped with Google’s rankings more than we forecast; both overlapped less.
That direction is the opposite of a problem for the thesis — it means the decoupling is, if anything, slightly stronger than we predicted, not weaker. The load-bearing criteria (blended, spread) are measured on the aggregate, and both cleared with room to spare. But it is a genuine miss on calibration: our AI Mode band was too high. The likely cause is that these buyer-intent SERPs are so affiliate-saturated that even AI Mode reaches past the top-10 to pull in vendor pages, Reddit threads, and niche comparison sites the organic ranking buries. We over-trusted the “built on Google’s index” mechanism.
How robust is this verdict to the missing ChatGPT data?
The real limit: ChatGPT completed only 3 of 8 queries this run (login/session walls on the rest — the recurring collection friction we have hit since the method post). Its 13.3% rests on three data points (CRM 20%, project management 0%, email 20%). So the fair question is: could full ChatGPT data flip the verdict?
We can bound it. The blended criterion is nearly unflippable — even if ChatGPT’s true overlap were as high as 50%, blended would be (53.6 + 36.4 + 50) / 3 ≈ 46.7%, still under the 50% ceiling. The spread criterion is more sensitive: it fails only if ChatGPT’s mean climbs above 28.6%. For that, the five uncollected queries would have to average above ~37.8% overlap — nearly triple the 13.3% we observed on the three that did complete. Given ChatGPT’s known lean toward brand-name recall over ranked-page retrieval, that is unlikely but not impossible. So we grade this a pass we are confident in on blended, and a pass we flag as data-limited on spread. We will re-run the full ChatGPT set and update if it moves.
What does this mean platform by platform? (the part you can act on)
The overlap number is directly actionable — it tells you how much of your AI-citation destiny classic Google ranking controls, per engine.
- Google AI Mode (53.6% overlap) — SEO is still half the battle, but only half. Ranking in the top-10 for your buyer query is the single biggest lever here; skip it and you forfeit the majority path. But ~47% of AI Mode’s citations come from outside the top-10, so a #1 ranking is necessary-not-sufficient. Do both: win the organic slot AND make sure the page carries the passage AI Mode wants to lift — a clean spec table, a direct “X vs Y” comparison, a dated verdict.
- Perplexity (36.4% overlap) — rank matters, citation play matters more. A third of Perplexity’s sources track Google’s rankings; two-thirds come from its own retrieval, heavy on review sites and community. Rank to stay eligible, but invest in being cited in the listicles and Reddit/forum threads Perplexity pulls — that is where the other two-thirds live.
- ChatGPT (13.3% overlap) — a nearly separate game. Rank barely predicted citation here. ChatGPT surfaced vendors that don’t sit in Google’s top-10 at all (for CRM it named Ecount, Freshworks, Pipedrive — none in the organic first page). Treat ChatGPT citation as its own discipline: brand-mention density, structured comparison content, and being named across the corpus the model draws on. Chasing a Google ranking to win ChatGPT is the least efficient move on the board.
The one-sentence policy: rank for Google AI Mode, get cited for ChatGPT, and do both for Perplexity. For a fuller per-engine playbook, see our per-engine GEO strategy checklist.
What are the limits of this verdict?
- One vertical of intent. Eight buyer-intent “best X” queries — the hardest case for the thesis, since these SERPs are the cleanest and most authoritative. A pass here is strong, but it describes commercial listicle search, not news, how-to, or navigational queries.
- A single logged-out snapshot. Engines and SERPs change weekly and personalize. This is one run; the same no-permanent-law caveat from the method post applies.
- ChatGPT n=3. The spread criterion clears on partial ChatGPT data. We flagged the exact flip threshold above and will refresh.
- Overlap is not causation. A wide gap shows citations diverge from rank — not that ranking is worthless. The positive tail everywhere is real: rank still buys the ticket.
What’s next?
This week measured how much AI citations overlap with rank. Next week we flip the question to why the gap exists. We will take the non-overlapping citations — the domains AI engines cite that do not rank in Google’s top-10 — and profile them: are they newer, more structured, more community-driven, more brand-named? If there is a repeatable signature of a “cited-but-not-ranked” page, that is the exact thing a GEO strategy should engineer. We will pre-register the predicted signature before we open the data, same as always.
Frequently asked questions
Does ranking #1 on Google still get you cited by AI?
It depends heavily on the engine. In our blind test of 8 buyer-intent queries, the share of AI-cited domains that also rank in Google’s top-10 was 53.6% on Google AI Mode, 36.4% on Perplexity, and 13.3% on ChatGPT. So on Google AI Mode a top ranking still lands you in the citations more than half the time; on ChatGPT it barely predicts citation at all. Rank buys the ticket, not the seat.
Is the “70% to under 20%” overlap collapse real?
Directionally yes — overlap is far below the old ~70% regime and is not uniform across engines. But “under 20%” specifically matches ChatGPT (13.3%), not the market. The blended overlap across three engines is 34.4%, and Google AI Mode sits at 53.6%. Quoting the sub-20% figure as the state of AI search cherry-picks the most decoupled engine.
How did you measure rank-to-citation overlap?
For each query we pulled Google’s top-10 organic domains and the distinct domains each AI engine cited. Overlap is the share of an engine’s cited domains that also appear in that query’s top-10, averaged across queries, excluding Google’s own properties and each engine’s self-citation. Blended is the mean of the three engines; spread is Google AI Mode minus ChatGPT. All bands were locked before collection.
What should I actually do differently per engine?
Rank for Google AI Mode (classic SEO controls ~half its citations), get cited for ChatGPT (brand mentions and structured comparisons matter far more than ranking), and do both for Perplexity. A #1 ranking is necessary-not-sufficient on AI Mode and nearly irrelevant on ChatGPT, so budget your effort accordingly.
Sources
- GeoParrot GEO Lab — blind run, 8 buyer-intent queries × 3 engines (Google AI Mode, Perplexity, ChatGPT web search), one logged-out snapshot, scored against Google top-10 organic. Full per-query scorecard in this experiment’s data.
- Does Ranking #1 on Google Still Get You Cited by AI? — Monday’s fact-check of the 70% → under 20% claim.
- Our pre-registered protocol — the frozen overlap bands and scoring rule.
- Per-engine GEO strategy checklist and Why there is no universal GEO strategy.

Leave a Reply