Quick answer: This is a pre-registration, not a result. On Monday we argued that ranking #1 on Google no longer buys you an AI citation — the overlap between top-ranked pages and AI-cited sources fell from roughly 70% to under 20% — but that the collapse is neither zero nor uniform across engines. Today we test that claim the hard way: a blind, pre-registered measurement of how much AI citations actually overlap with Google’s rankings. Two rival models compete. The rank-mirror model says AI answers are essentially a re-skin of Google’s top-10, so overlap should be high and uniform across engines. Our rank-decoupled model says citations diverge from rank — and the divergence is uneven per engine, because Google AI Mode is built on Google’s own index while ChatGPT and Perplexity crawl their own way. We locked 8 buyer-intent “best [category] 2026” queries and wrote the predicted overlap bands down before running anything: Google AI Mode 55–75%, Perplexity 25–45%, ChatGPT 15–35%, blended ≤50%, with an engine spread of ≥25 points. Wednesday we collect blind, Thursday we score, Friday we rule. The rank-decoupled thesis passes only if blended overlap lands at or below 50% and the AI-Mode-minus-ChatGPT spread is at least 25 points; we retract if overlap tops 55% and the engines move together.
This is Tuesday in this week’s GEO Lab arc, and it follows the same discipline we used last week. Monday’s fact-check made a claim — that rank is “a ticket, not the seat.” A claim that only describes the third-party numbers we quoted is cheap. The honest follow-through is to make it predict our own first-party measurement, with the numbers written down first. That’s a pre-registration: a public, timestamped commitment that removes our freedom to reinterpret the result after it lands.
What exactly are we measuring — and why overlap, not rank correlation?
Monday leaned on other people’s aggregate statistics (a 70%→20% overlap collapse, an r ≈ 0.18 correlation). This week we generate the number ourselves, on a metric a real buyer can feel: overlap. For a given query, we take the set of sources an AI engine actually cites in its answer and ask what share of them also rank in Google’s top-10 organic results for the same query. High overlap means the AI is basically reading Google’s rankings back to you. Low overlap means it is sourcing somewhere Google’s rankings don’t reach. We chose overlap over a correlation coefficient on purpose: correlation compresses the whole story into one abstract number, while overlap is directly actionable — it tells a page owner whether ranking on Google is the path to the citation, or a different game entirely.
What are the two rival models, stated as mechanical predictions?
We compress the whole week into two worldviews that make different, falsifiable predictions about the same data:
- Rank-mirror model (the null we’re testing against). AI search is a convenience layer on top of classic search: it retrieves the same top-ranked pages and summarizes them. If that’s true, the sources it cites should overlap heavily with Google’s top-10 (we predict >50%), the overlap should be similar across all three engines, and cited pages should cluster in Google’s top-3.
- Rank-decoupled model (our hypothesis). AI citations are drawn from a different distribution than Google’s blue links — community threads, docs, vendor pages, and mid-tail sources that rank poorly or not at all. And critically, the decoupling is uneven per engine: Google AI Mode sits on Google’s index and stays closest to it, while ChatGPT and Perplexity — with their own retrieval and a well-documented lean toward Reddit and forums — pull furthest away. This is the no-universal-GEO-strategy pattern applied to rank itself.
When both models predict the same thing, the data tells us nothing. The test has teeth only where they diverge — and here they diverge on two axes at once: the level of overlap (high vs. low) and the spread across engines (uniform vs. uneven). We score both.
Which 8 queries did we freeze — and why buyer-intent?
We fixed the query set before looking at any AI or Google output. The rule: high-commercial-intent “best [category]” searches, the exact kind where Google’s top-10 is dominated by affiliate listicles (PCMag, Forbes Advisor, TechRadar, G2). Those SERPs are the hardest case for our thesis — if AI just mirrored Google, buyer-intent queries with their tidy listicle rankings are where the overlap should be highest. Betting on decoupling here is betting against the odds on purpose.
| # | Frozen query (asked verbatim, 2026) | Why it’s a hard case for us |
|---|---|---|
| 1 | best CRM for small business | Listicle-saturated SERP (HubSpot/Salesforce/Zoho affiliate pages) |
| 2 | best project management software | Asana/Monday/ClickUp all rank and advertise heavily |
| 3 | best email marketing platform | Mailchimp vs. Klaviyo vs. Brevo — pure affiliate territory |
| 4 | best web hosting for WordPress | Among the most SEO-optimized SERPs on the web |
| 5 | best VPN | Notoriously affiliate-driven; top-10 is nearly all listicles |
| 6 | best password manager | 1Password/Bitwarden dominate both rankings and reviews |
| 7 | best accounting software for freelancers | QuickBooks/FreshBooks/Wave affiliate pages own the top-10 |
| 8 | best AI writing tool | Fast-moving category where community sources may diverge most |
“Blind” here means we have not run any of these eight on ChatGPT, Perplexity, or Google AI Mode, and we have not pulled the Google top-10 for them yet. We know these categories are affiliate-heavy; what we do not know is which specific domains each engine will cite this week, or how far those domains sit from Google’s rankings. That gap is exactly what Wednesday measures.
The pre-registered overlap bands (locked)
This table is the frozen artifact. Each band was written before any collection, so Thursday’s write-up cannot move the goalposts. Overlap = the share of an engine’s distinct cited domains (per query) that also appear in that query’s Google top-10 organic, averaged across the 8 queries.
| Engine | Predicted overlap band | Why this band |
|---|---|---|
| Google AI Mode | 55–75% | Generated on top of Google’s own index — should stay closest to the rankings (highest by design) |
| Perplexity | 25–45% | Own retrieval and citation UI; moderate Google divergence, lighter community skew than ChatGPT |
| ChatGPT (web search on) | 15–35% | Heaviest documented lean toward Reddit/forums/mid-tail — furthest from Google’s rankings (lowest by design) |
| Blended (all 3) | ≤ 50% | Weighted overall — the headline “rank ≠ citation” figure |
| Engine spread (AI Mode − ChatGPT) | ≥ 25 points | The unevenness that separates our model from the rank-mirror null |
Two numbers carry the test. The blended level (≤50%) tests whether rank and citation have decoupled at all. The engine spread (≥25 points) tests whether the decoupling is uneven — the fingerprint of our model versus a uniform mirror. We also log a secondary position test: of the citations that do overlap with Google, we predict ≥50% come from Google’s top-3 for AI Mode, but no top-3 concentration for ChatGPT (its overlapping cites should scatter across ranks 1–10).
How will Wednesday’s blind collection work?
To keep the outcome uncontaminated, the protocol is fixed here, in advance:
- One neutral query per category. The eight strings above, asked verbatim, logged out, US locale, single run — no follow-ups, no re-rolls.
- Two captures per query. (a) The AI answer’s cited sources on each of the three engines — ChatGPT (web search on), Perplexity, Google AI Mode — deduplicated to distinct domains. (b) Google’s top-10 organic domains for the same query, excluding ads, the AI Overview block, People-Also-Ask, and Google’s own properties.
- Overlap defined now. A cited domain “overlaps” if it appears anywhere in that query’s Google top-10 organic. We match on registrable domain (so a deep TechRadar article counts against a TechRadar top-10 listing). Google’s own domains and each engine’s self-citation are excluded from both sides.
- Scoring rule. Per engine, overlap = mean across the 8 queries of (distinct overlapping cited domains ÷ distinct cited domains). Blended = mean of the three engine scores. Spread = AI Mode minus ChatGPT.
- No touching the bands. Wednesday’s collector records raw citations and rankings only; it never sees or edits the predicted-band table above.
What result would prove the rank-decoupled model wrong?
The pass/fail line is committed now, so Thursday’s numbers decide it, not our narration:
- Blended overlap ≤ 50% AND engine spread ≥ 25 points → the rank-decoupled model is confirmed for buyer-intent queries: rank and citation have separated, and the separation is engine-specific. Monday’s “ticket, not the seat” claim graduates from a quoted third-party stat to our own first-party measurement.
- Only one of the two holds → inconclusive. Reported as not passing. (Low overlap but uniform across engines would mean AI diverges from Google everywhere equally — decoupled, but not the per-engine story we predicted. High overlap but a wide spread would mean rank still mostly predicts, with one engine as an outlier.)
- Blended overlap > 55% AND spread < 15 points → we retract. That is the rank-mirror model winning: AI is largely re-surfacing Google’s top-10, uniformly. We would say plainly, in this same place, that Monday’s decoupling framing was overstated for buyer-intent search.
Notice the retract condition is a real possibility, not a strawman. Buyer-intent listicle SERPs are precisely where Google’s rankings are cleanest and most authoritative, so a high, uniform overlap is a live outcome we’ve committed to accept if it lands.
What are the limits of even a passing result?
A clean pass would be genuine evidence, but we won’t oversell it:
- Eight queries, one vertical of intent. All eight are commercial “best X” searches. A pass says rank and citation decoupled for buyer-intent, not for informational or navigational queries — where the overlap could look very different.
- Domain-level matching. We match registrable domains, not exact URLs. That’s deliberate (a citation to any TechRadar page should count against TechRadar’s ranking presence), but it inflates overlap slightly versus strict URL matching. We’ll report both if the gap is material.
- Snapshot and personalization. Engines and SERPs change weekly and personalize; this is a single logged-out run. A pass describes this snapshot, not a permanent law — the same no-universal-strategy caveat applies.
- Overlap is not causation. Even a wide gap shows citations diverge from rank — not that ranking is worthless. Low overlap with a still-positive tail (Monday’s r ≈ 0.18) is entirely consistent: rank buys the ticket, the citation is a separate draw. We keep the causal claim out of scope, as we did in our backlinks analysis.
That’s the honest frame for the week. We took the standard we apply to every third-party GEO stat and pointed it at Monday’s own claim — with our own data, our own eight queries, and the pass line fixed in advance. If the overlap lands low and the engines spread apart, the rank-decoupled model earns the word “model.” If Google’s rankings turn out to still drive the citations, you’ll read that here on Friday. You can see our full first-party scorecard in the 2026 GEO Benchmark.
Frequently asked questions
How is this different from Monday’s article?
Monday was a fact-check of other people’s aggregate numbers — the reported 70%→20% overlap collapse and the r ≈ 0.18 correlation. Today is our own blind experiment: eight frozen queries, three engines, and overlap bands we locked before collecting a single data point. Monday asked “is the claim true?”; this week measures it first-hand.
Why measure overlap instead of just checking if the #1 result gets cited?
Because a single-position check is noisy and easy to cherry-pick. Overlap across a query’s whole top-10 and a whole citation set is a stable, buyer-relevant number: it answers “if I rank on Google for this, how often does that actually put me in the AI’s sources?” — which is the question a page owner is really asking.
Isn’t picking the queries yourselves a way to rig the result?
We stacked the deck against ourselves on purpose. All eight are affiliate-heavy buyer-intent SERPs — the cleanest, most SEO-optimized rankings on the web, where a rank-mirror would show the highest overlap. If decoupling shows up even here, it’s a conservative result. The frozen query list and pre-committed bands are the guardrail against reinterpreting a near-miss later.
Why do you predict Google AI Mode will overlap most?
Because AI Mode generates its answer on top of Google’s own search index, so its sources should track the rankings more closely than engines with independent retrieval. If AI Mode’s overlap is not the highest of the three, that alone would be a surprising finding worth flagging — and it’s why the engine spread, not just the blended level, is a load-bearing prediction.
How can I run a version of this on my own category?
Pick one “best [your category] 2026” query. Pull Google’s top-10 organic domains. Then ask the same query on ChatGPT, Perplexity, and Google AI Mode and list the domains each one cites. Count how many cited domains also sit in Google’s top-10 — that ratio is your overlap. If it’s low, ranking on Google is a ticket, not the seat, and you need a separate AI-citation play on top of classic SEO.
Sources
- GeoParrot GEO Lab — Does ranking #1 on Google still get you cited by AI? (this week’s Monday fact-check, the claim under test).
- GeoParrot GEO Lab — last week’s pre-registration method (the same blind, locked-prediction discipline).
- GeoParrot GEO Lab first-party data — cross-engine citation overlap results, does AI search use backlinks, why there’s no universal GEO strategy, and the 2026 GEO Benchmark.

Leave a Reply