We Bet Rewording Wouldn’t Change Who AI Cites. On Google AI Mode, It Swapped ~86% of the Sources.

Quick answer: Our load-bearing prediction (P1) failed. We predicted that rewording the same buyer intent would leave AI’s cited sources roughly as stable as re-running the identical query — that phrasing is a doorbell, not a gatekeeper. Instead, on Google AI Mode, two different phrasings of the same intent shared only 0.14 of their cited sources (Jaccard), while running the identical query twice shared 0.53. That is a 0.40 gap — the pre-registered design allowed at most 0.15. Rewording moved the citation panel roughly 4× more than run-to-run noise: about 86% of the pooled sources differed between phrasings. One prediction held (P2: long queries did not skew toward niche sources), one failed alongside P1 (P5: only 2 of 8 topics kept any source across every phrasing), and one (P4, engine dependence) we could not test because the ChatGPT arm was blocked mid-run. This is the results post for Tuesday’s pre-registered design; nothing below was reframed after the numbers came in.

Experiment results — published September 24, 2026, in the GEO Lab. We locked the topics, the three phrasing tiers, the noise-floor control, and five falsifiable predictions on Tuesday, collected citations Wednesday, and computed every overlap before writing a word of this post.

What did we actually measure?

We took 8 buyer-intent topics — project management, CRM, robot vacuum, noise-cancelling headphones, VPN, password manager, standing desk, email marketing — and wrote each one in three ways: a bare keyword (“crm software”), a mid-length qualifier (“best crm software for solo consultants”), and a long conversational question (“which crm is easiest to set up for a solo consultant who hates complicated tools?”). Same intent, three phrasings. Then we added the piece that makes this a real experiment: we ran the mid-length phrasing a second, identical time — the noise floor. That gives us a baseline for how much AI’s citation panel wobbles on its own, before phrasing is even in the picture.

For every query we pulled the distinct sources Google AI Mode cited, then measured overlap between source sets with the Jaccard index (shared sources ÷ total unique sources; 1.0 = identical panels, 0 = no shared source). The comparison that decides the whole experiment is simple: does rewording the query change the cited sources more than the query changes on its own when you just run it twice? If phrasing is a doorbell, cross-phrasing overlap should sit right up near the noise floor. If it is a gatekeeper, cross-phrasing overlap should collapse well below it.

Did rewording change who AI cites, like we predicted it wouldn’t?

It did — decisively, and in the opposite direction from our bet. Averaged across the 8 topics, the noise floor (same query, run twice) came in at 0.53 Jaccard: re-running an identical query keeps a little over half the cited sources. But cross-phrasing overlap — the average of the three phrasing pairs per topic — was just 0.14. Rewording the same intent kept only about one source in seven. Put the other way: roughly 86% of the pooled sources differed between two phrasings of the same buyer question.

Grouped bar chart across 8 topics: noise-floor overlap (same query re-run) averages 0.53 while cross-phrasing overlap averages 0.14; rewording moves citations about 4x more than re-running
Every topic except headphones shows the same shape: re-running the identical query (navy) holds far more of the citation panel than rewording the same intent (coral). Overlap = Jaccard on the cited-source sets.

The one exception is instructive. Headphones is the only topic where cross-phrasing overlap (0.30) beat its own noise floor (0.20) — but that is because its re-runs were themselves unusually unstable, and because RTINGS.com anchored every phrasing. When a topic has one dominant, universally trusted reviewer, phrasing matters less. That is the shape of a genuine exception, not a measurement glitch — and it previews the practical takeaway below.

Is a 0.14 overlap actually below the noise, or just measurement fuzz?

This is exactly why the noise floor was pre-registered. Google AI Mode obfuscates its citation links, so we match sources on publisher name, not clean domain — a noisier signal. But the noise floor is measured with the same name-matching method, so that fuzziness is baked equally into both bars. The 0.53-vs-0.14 gap survives it. P1 locked a threshold before collection: cross-phrasing overlap had to land within 0.15 Jaccard of the noise floor to count phrasing as “invariant.” The actual gap was 0.40 — more than double the allowed slack.

Bar chart: same-query re-run overlap 0.53 versus reworded-same-intent overlap 0.14, with the pre-registered P1 threshold at 0.38; cross-phrasing falls far below the threshold, a 0.40 gap
P1 required cross-phrasing overlap to stay above 0.38 (noise floor minus the 0.15 allowance). It landed at 0.14 — a 0.40 gap. On this engine, phrasing is closer to a gatekeeper than a doorbell.

Here is the full per-topic sheet, so you can see the gap is broad, not carried by one outlier:

Topic Noise floor (same query) Cross-phrasing (reworded) Gap Phrasing-invariant source?
Project management 0.83 0.25 +0.58 YouTube
CRM 0.38 0.10 +0.28 — none
Robot vacuum 0.57 0.03 +0.54 — none
Noise-cancelling headphones 0.20 0.30 −0.10 RTINGS.com
VPN 0.50 0.16 +0.34 — none
Password manager 0.50 0.14 +0.36 — none
Standing desk 0.60 0.05 +0.55 — none
Email marketing 0.67 0.06 +0.61 — none
Mean 0.53 0.14 +0.40 2 of 8 topics
Seven of eight topics show a positive gap (noise floor above cross-phrasing). Gap = noise floor − cross-phrasing overlap.

Did long, conversational queries drift toward niche sources? (P2)

No — and this is the prediction that held. A common GEO worry is that long-tail conversational queries pull AI away from big authority sites toward obscure blogs. We measured authority share: the fraction of a tier’s citations going to publishers that showed up across two or more queries (a demand/authority proxy). P2 predicted the drop from short to long queries would stay under 15 points. It did not drop at all — authority share actually rose, from 46% on bare keywords to 55% on long conversational questions. Long queries did not reach for niche sources; if anything they leaned harder on repeat-cited authorities.

Read P1 and P2 together and the real finding sharpens: phrasing reshuffles which authorities get cited, not authority-for-niche. The panel stays roughly as brand-heavy — but the specific brands rotate out with the wording.

Does any source survive every phrasing? (P5)

Mostly not. P5 predicted that at least 60% of topics would keep a “phrasing-invariant core” — a source cited across all three phrasings. Only 2 of 8 (25%) did: YouTube for project management and RTINGS.com for headphones. Six of eight topics had no source that survived all three wordings. P5 failed — but the two survivors are the tell: an invariant anchor exists only where one destination utterly dominates a category (video walkthroughs for PM tools; RTINGS for audio gear). Everywhere else, the citation panel is phrasing-contingent all the way down.

What happened to the ChatGPT arm? (P4)

We owe you full transparency here, because the pre-registered design named two engines and we are publishing one. When we collected on Wednesday, ChatGPT’s logged-out web search had been gated behind login — the same queries that returned live review-site citations in our September 17 run now returned no search at all (no “searched the web” banner, no source panel, just model-memory prose). That is an OpenAI platform change, not a collection bug, and we caught it mid-run.

We chose not to paper over it by logging in. A logged-in session brings personalization and memory into a phrasing experiment — the one confound that would quietly contaminate the exact signal we are measuring. So P4 (engine dependence) is deferred, not tested. This week’s result stands as a clean, single-engine finding on Google AI Mode; we cannot yet say whether ChatGPT behaves the same. That is the honest state of the data, and it is a live example of why single-snapshot GEO measurement is fragile: the ground moves under you between runs.

What does this mean if you are trying to get cited?

Three practical reads, held to what one engine’s snapshot can actually support:

  • Stop optimizing for “the” query. There is no single citation panel for a buyer intent — there is a different one for nearly every way a person can phrase it. Winning one phrasing tells you almost nothing about the next. This is the citation-side twin of query fan-out: one intent explodes into many retrievals.
  • Breadth beats a single perfect page. If ~86% of sources rotate with phrasing, the way to show up often is to be citable across many phrasings of the intent — comparison angles, persona-specific takes, constraint-specific answers — not one keyword-perfect landing page.
  • Category dominance is the only true invariant. The sources that survived every phrasing (YouTube, RTINGS) are the ones that own their category’s trust. Short of that, aim to be one of the rotating authorities — authority share stayed high, so brand demand still buys you a seat, just not a fixed one.

The honest limits

  • One engine. Google AI Mode only; the ChatGPT arm was blocked. We cannot claim this generalizes across engines — that is precisely the untested P4.
  • Publisher names, not domains. AI Mode obfuscates links, so overlap is computed on publisher name. The noise-floor control absorbs this (same method, both bars), but exact-domain matching could shift individual numbers.
  • Single snapshot, logged out. One collection pass on one day. The noise floor tells us within-day wobble is real (0.53, not 1.0); across days it would be larger.
  • Small panels. AI Mode cited 2–7 sources per query, so single-source swaps move Jaccard a lot. The pattern is consistent across 8 topics, but each topic is a small sample.

None of these soften the headline: on the engine we could measure, rewording the same intent cleared the citation panel far below what re-running the identical query does. Friday’s verdict post will hold all five predictions up against the numbers, score the “doorbell vs gatekeeper” bet, and set up next week’s question — whether a single dominant source (RTINGS-style) is the only thing that survives phrasing, across engines.

Frequently asked questions

Does rewording a query really change which sources AI cites?

On Google AI Mode, yes. Across 8 buyer-intent topics, two different phrasings of the same intent shared only 0.14 of their cited sources (Jaccard), versus 0.53 for the identical query run twice. Roughly 86% of the pooled sources differed between phrasings — about 4× more turnover than run-to-run noise.

What is a noise floor and why does it matter here?

The noise floor is how much the citation panel changes when you run the exact same query twice, changing nothing. It sets the bar for a real phrasing effect: cross-phrasing overlap must fall clearly below it. Ours did — 0.14 versus a 0.53 floor — so the effect is signal, not measurement wobble.

Do long conversational queries pull AI toward niche sources?

Not in this experiment. Authority share — the fraction of citations going to repeat-cited, high-demand publishers — actually rose from 46% on bare keywords to 55% on long questions. Phrasing rotates which authorities get cited; it does not trade authority for obscurity.

Why only Google AI Mode and not ChatGPT?

The design named both, but ChatGPT gated its logged-out web search behind login between our September 17 run and this one. Logging in would inject personalization into a phrasing test, so we deferred the ChatGPT arm rather than contaminate the signal. Engine dependence (P4) is untested this week.

What should I do to get cited if phrasing is this volatile?

Be citable across many phrasings of an intent rather than optimizing one keyword-perfect page: comparison angles, persona and constraint-specific answers. The only sources that survived every phrasing were category-dominant destinations (YouTube for project management, RTINGS for headphones), so category trust is the real invariant.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

🦜 Follow GeoParrot: YouTubeX