Quick answer: Last week’s experiment turned up a number that reframes all of GEO: when we re-ran the identical query on Google AI Mode a second time, the two answers shared only 53% of their cited domains (Jaccard 0.53). Half the citation panel rotated with nothing changed — same words, same session, minutes apart. That is the noise floor, and it raises the question that actually matters if you’re spending money to get cited: if a panel wobbles that much on its own, does a citation you earn today survive to next week, or does it decay away on its own? This week we test it directly. The design: a fixed set of 8 buyer-intent queries, snapshotted once a day for several consecutive days on Google AI Mode, measuring how much of Day 0’s cited-domain set is still there on each later day. We compare that temporal overlap against last week’s same-session noise floor — because a citation can only be called “durable” if it survives time better than it survives a plain re-run. Our locked bet: AI citations are rented, not owned — but the rent has a floor. A small authority-weighted core stays put every day while the long tail turns over. Five numbered predictions are pinned below before we collect anything. Collection starts today.
Experiment design — published September 29, 2026, in the GEO Lab. This is a pre-registration: the query set, the snapshot schedule, the metrics, and five falsifiable predictions go out now; we collect the daily snapshots this week and open the data later. We are dogfooding our own volatility measurement — the same discipline that stopped us fooling ourselves last week, when the data landed in the opposite direction from our bet.
What claim are we testing this week?
The hypothesis, stated so it can lose: a citation in an AI answer is a durable asset — once an engine decides your page belongs in the panel for a query, it keeps citing you day after day, the way a #1 organic ranking used to sit still for weeks. On that story, GEO is an investment that compounds: earn the citation once, keep the traffic and the authority signal for the foreseeable future.
The counter-hypothesis, which last week’s noise floor leans hard toward: an AI citation is a rented spot on a shelf that gets re-stocked constantly. If re-asking the identical question already swaps half the sources within minutes, then over days the panel may drift so much that “being cited on Tuesday” tells you almost nothing about Thursday. On that story, GEO isn’t a one-time win you bank — it’s a position you have to keep re-earning, and any tool that reports “you were cited” from a single snapshot is selling you a lottery ticket as if it were a deed.
Both are plausible, and the whole GEO-tracking industry quietly assumes the first — visibility dashboards show a citation as a status you hold, not a coin flip you re-run. This experiment is built to make the data pick a side, and to separate two things people conflate: how much a panel wobbles on its own (the noise floor) versus how much it actually drifts over time (the thing you’d call decay).
What exactly gets collected — the queries and the schedule?
Eight commercial, buyer-intent queries — the ones where a citation carries real money — held word-for-word identical across every snapshot, so the only variable is time. These are the same intents we’ve tracked all series, phrased as the mid-length natural queries a buyer actually types:
| Intent | Exact query (frozen across all snapshots) |
|---|---|
| Project management tool | “best project management software for small teams” |
| CRM | “best crm software for solo consultants” |
| Robot vacuum | “best robot vacuum for small apartments” |
| Noise-cancelling headphones | “best noise cancelling headphones for the office” |
| VPN | “best vpn for streaming while traveling” |
| Password manager | “best password manager for families” |
| Standing desk | “best standing desk for a home office” |
| Email marketing platform | “best email marketing platform for small ecommerce” |
We run a single engine — Google AI Mode — on purpose. Last week’s most interpretable numbers came from holding one engine fixed, because the two engines retrieve so differently that blending them hides the answer. A temporal-stability test needs that same precision: to say “this domain dropped out on Day 2” you need the measurement floor to be as quiet as possible. Perplexity sits out again — it rate-limits hard mid-collection, which would confound a decay signal with a collection artifact, the one mistake this design exists to avoid.
For each snapshot and query we capture the set of distinct domains AI Mode cites. Then, for every later day N, we compare Day N‘s set back to Day 0’s set. This is a pilot window of a handful of consecutive days, not a month — we’re explicit about that, and a longer multi-week follow-up is already scoped. But even a few days is enough to answer the load-bearing question: does time erode a citation panel faster than a plain re-run does?
How do we tell real decay from the panel’s own wobble?
This is the crux, and it’s the step that separates a measurement from a vibe. If Day 2’s citations differ from Day 0’s, some of that difference is genuine drift over time and some is just the same run-to-run randomness we already measured last week. You cannot call anything “decay” until you know how much the panel moves when nothing changes. Luckily, we already have that number.
The same-session noise floor from last week was 0.53 — re-asking the identical query, seconds apart, kept 53% of the cited domains. That’s our benchmark. If the day-to-day overlap sits right around 0.53, then time adds nothing beyond the panel’s built-in wobble: your citation is exactly as (un)stable on Day 3 as it was on a same-minute re-run, and “decay over time” is a myth — the panel was always this loose. If the day-to-day overlap sits well below 0.53, time is doing real damage on top of the noise, and a citation genuinely erodes. We measure “how much the set moved” with the Jaccard overlap — shared domains divided by combined domains, where 1.0 means an identical panel and 0.0 means no domain in common. Four numbers do the work:
- Temporal overlap — mean Jaccard between Day 0 and each later day’s cited-domain set, per query. The headline number: how stable citations are as days pass.
- Domain survival rate — of the domains cited on Day 0, the fraction still cited on the final day. A person-level view: if you got cited Monday, what are your odds of still being there Friday?
- Invariant core — the domains cited on every snapshot for a query. The durable backbone, if one exists at all.
- Authority-weighted survival — survival rate split by tier, using the same high-brand-demand concentration math as our 15-domains study. Tells us whether the core that survives is the big authority sites or something else.
The comparison that decides the experiment is a single inequality: is the day-to-day temporal overlap lower than the 0.53 same-session floor, or not? Everything else — the survival curve, the core, the authority split — explains the shape of the answer. But that one comparison is what turns “AI citations feel volatile” into a measured claim you can act on.
What are the five predictions we’re locking before collection?
Pinned in public, before a single snapshot is scored, so we can be caught out if the data disagrees. Our overall bet: rented, not owned — but the rent has a floor.
- P1 (load-bearing) — Time adds churn, but not total collapse. Day 0 → final-day temporal overlap will land below the 0.53 same-session noise floor (real decay exists), yet stay above 0.20 (the panel doesn’t fully reshuffle). If temporal overlap sits at or above 0.53, decay is a myth and we’re wrong.
- P2 — An invariant core exists. For the majority of the 8 queries, at least one domain is cited on every snapshot. There is a backbone, even if it’s thin.
- P3 — The core is authority-weighted. High-brand-demand / authority domains survive Day 0 → final day at a meaningfully higher rate than long-tail domains. What you can “hold” is the top of the panel, not the tail.
- P4 — Decay is monotonic. Overlap shrinks as the lag grows: Day 0 → Day 1 overlap is higher than Day 0 → final day. Drift accumulates rather than bouncing.
- P5 — UGC anchors persist. Reddit and YouTube — which were the invariant sources across phrasings last week — are among the most persistent domains over time too, showing up in the invariant core disproportionately often.
Note how these can conflict, which is what makes them a real test rather than a horoscope: P1 says the panel churns a lot, while P2, P3 and P5 say a stubborn core resists that churn. If both hold, the picture is a stable authority spine with a constantly-rotating long tail — the “rented with a floor” model. If P1 holds but P2/P3/P5 fail, citations are pure lottery and no one is safe. If P1 fails, the industry’s “citation as durable asset” assumption survives and we eat our bet in public. Any of those is a publishable result.
Why does this matter for how you spend on GEO?
Because it changes what a “win” is worth. If citations are durable, the GEO playbook is classic SEO: invest to earn the spot, then defend it, and model the traffic as an annuity. If citations are rented with a floor, the math flips — the ROI lives entirely in whether you’re authority-tier enough to sit in the persistent core, and everyone below that line is buying a spot that evaporates, no matter how perfectly they optimized the page. That also indicts single-snapshot tracking: a dashboard that checked once and told you “you’re cited” may have caught you on a good coin flip. The honest metric isn’t “were you cited” — it’s “what’s your survival rate,” and almost nobody reports that. This is also why we keep flagging that citation volatility is a different beast from ranking volatility: rankings move positions, but citation panels swap their entire cast.
Collection runs this week; the daily snapshots and the survival curve go out in the results post, with the raw sets attached so you can audit every domain that dropped. If our load-bearing P1 is right, the takeaway will be uncomfortable and specific: stop treating an AI citation as something you own, and start measuring whether you can keep it.
Frequently asked questions
What’s the difference between the “noise floor” and “decay”?
The noise floor is how much a citation panel changes when you ask the identical question twice in the same session — last week that was 0.53 Jaccard on Google AI Mode, meaning 47% of the panel swapped with nothing changed. Decay is how much the panel changes over days. Decay only counts as real if the day-to-day overlap is lower than the same-session noise floor; otherwise the panel was simply always that loose, and time added nothing.
Why only Google AI Mode and not ChatGPT or Perplexity?
Precision. Measuring small day-to-day movement requires the quietest possible measurement floor, and mixing engines that retrieve differently would blur the signal. Perplexity also rate-limits mid-collection, which would masquerade as decay. Holding one engine fixed is the same discipline that gave us last week’s clearest numbers.
Isn’t a few days too short to call it “decay over time”?
For a definitive lifespan, yes — that’s why we call this a pilot and have a multi-week follow-up scoped. But a few days is enough to answer the load-bearing question: does time erode a citation panel faster than a plain re-run does? If overlap already drops below the same-session floor within days, decay is real and worth measuring at longer horizons.
What would prove you wrong?
If Day 0 → final-day overlap lands at or above 0.53 (the same-session noise floor), decay is a myth and the “citation as durable asset” view wins. If no invariant core appears for most queries, the “rent has a floor” half of our bet fails and citations look like pure lottery. Both outcomes get published exactly as they land.
How can I re-run this myself?
The eight frozen queries are listed above, and the results post will ship every daily snapshot as raw cited-domain sets plus the scoring script. Run the same eight strings on Google AI Mode logged-out, once a day, record the cited domains, and compute the Jaccard overlap between your first day and each later day. If your survival rates differ from ours, that itself is a finding worth sharing.

Leave a Reply