Quick answer: An earned narrative beats an incumbent when it’s pointed at a distinct axis the incumbent structurally can’t claim — not when the incumbent happens to be small. The incumbency moat is a geometry gate, not a size threshold. On our frozen 6-category roster, with zero engines re-queried, the incumbent’s size advantage drew no separating line: the largest size gap in the study — Mullvad’s 176 reviews against ExpressVPN’s 28,215, a 160x gap — got crossed, while a much smaller 2.2x gap in accounting held. There is no review-count or age threshold beyond which a narrative can’t win. What did separate the two crossed categories from the four held ones, 6 of 6, was the shape of the challenger’s narrative: every challenger that owned a distinct, orthogonal axis the incumbent can’t stand on (Mullvad → privacy/no-logs, HubSpot → free-CRM + app-marketplace ecosystem) crossed the moat; every challenger selling a “better version” of the incumbent’s own axis (Xero’s “better reporting,” Freshdesk’s “omnichannel too”) stayed held, size gap notwithstanding. So the market’s “AI levels the playing field” is half true — the door is open, but the room is won on axis geometry. The honest catch: this is retrodiction on categories whose outcomes we already knew, n=6 with only 2 crossed, so this is a strong working rule, not a proven law — and we test it blind next week. Below: the ruling, the scorecard, a decision test for your category, a per-engine playbook, and the limits.
This is Friday’s verdict, closing this week’s GEO Lab arc. Monday we fact-checked the loudest optimistic GEO claim of the month — “AI search levels the playing field for challengers” — and argued it’s half true: two classic incumbent moats (domain authority, backlinks) genuinely dissolved, but a third survived, and earned narrative crosses it only sometimes. Tuesday we pre-registered a method pitting two rival explanations against each other — a magnitude model (the moat is the incumbent’s size) versus a geometry model (the moat is the shape of the axis you compete on) — and locked four predictions before scoring a single category. Wednesday we coded the frozen roster with zero engine re-queries. Thursday we published the results: four for four. Today we rule on the week’s question and turn it into a playbook.
What exactly were we ruling on?
One question threaded through four posts: when does a challenger’s earned narrative actually beat an incumbent in AI search — and where does the incumbency moat hold? Last week we established that narrative ownership is earnable, not locked in category history — a challenger can deliberately author the language of a category. But we also found it’s necessary, not sufficient: in half the categories where a challenger owned the narrative, the AI still picked the deep incumbent. That left the exact question this week had to settle: what decides which side of the line you land on?
The two candidate answers make opposite predictions, so the test is clean:
- If the moat is magnitude (size), then challengers should cross when the incumbent’s size advantage is small and stay held when it’s large — a monotone threshold you could draw on a size axis.
- If the moat is geometry (axis shape), then size should be irrelevant, and what should separate crossed from held is whether the challenger owns a distinct axis the incumbent can’t claim, or merely a better version of the incumbent’s own.
We couldn’t run a live A/B — you can’t randomize which brand “owns privacy.” So this is a natural experiment on the frozen 18-brand, 6-category roster from the earlier weeks of this thread, coded blind to the size numbers, with the engines never re-queried (re-querying mid-experiment would have broken the pre-registration). Here’s how the two models scored.
Did the incumbent’s size decide it? ❌ No — the moat is not a wall of scale
Refuted, and not by a little. If size were the moat, the two crossed categories would cluster at small size gaps and the held ones at large gaps. They don’t overlap a little — they invert. The crossed categories span the entire range, from marketing automation at 1.1x to VPN at 160x, while the two genuinely-contested held categories sit right in the middle: help desk at 1.85x and accounting at 2.22x. Held accounting has a bigger size gap than crossed marketing automation. There is no cut you can draw on the size axis that puts the crossed categories on one side and the held ones on the other.

The single most decisive row is VPN. Mullvad carries 176 independent reviews against ExpressVPN’s 28,215 — the largest size gap in the entire study, the one the magnitude model is most confident should hold — and it crossed anyway. The engines’ #1 VPN is Mullvad. Magnitude says ExpressVPN by 160x; the engines pick the brand with 160x fewer reviews. This is the same reversal that has recurred all thread: a brand can be last on volume-style metrics yet first on what it owns. Verdict on the size question: there is no incumbency-size threshold. The moat is a gate, not a wall.
So what actually decides it? ✅ The geometry of the challenger’s axis, 6 of 6
Confirmed, cleanly, on this set. Coding each category’s owned attribute blind to the size numbers into two buckets — a distinct axis (an orthogonal trade-off the incumbent structurally can’t claim) versus a same axis (a “better version” of a dimension the incumbent already stakes) — classified every category correctly. Both distinct-axis challengers crossed; all four same-axis categories held. Zero misclassifications.
| Category | Challenger’s axis | Axis geometry | Size gap | Outcome |
|---|---|---|---|---|
| VPN | Mullvad owns privacy / audited no-logs / anonymity | Distinct | 160x | ✅ Crossed |
| Marketing automation | HubSpot owns free-CRM + app-marketplace ecosystem | Distinct | 1.1x | ✅ Crossed |
| Accounting | Xero sells “better reporting / visibility” | Same | 2.2x | ❌ Held (QuickBooks) |
| Help desk | Freshdesk sells “omnichannel too” | Same | 1.85x | ❌ Held (Zendesk) |
| Live chat | no distinct-axis challenger attempted | Same | — | ❌ Held (LiveChat) |
| Survey | no distinct-axis challenger attempted | Same | — | ❌ Held (SurveyMonkey) |
Read the split as one rule: a “better version” of the incumbent’s own axis reinforces the incumbent’s frame instead of escaping it. Xero’s “better reporting” is a comparative on QuickBooks’ axis — reporting is QuickBooks’ core territory — so the engines keep QuickBooks. Freshdesk’s “omnichannel” is Zendesk’s territory too (Zendesk Suite in 2018 preceded Freshdesk’s Omniroute in 2019), so it held. Mullvad, by contrast, didn’t offer a faster ExpressVPN; it answered a different question — who protects your privacy? — that ExpressVPN’s speed-and-streaming positioning structurally can’t claim. Two of the four held cases (live chat, survey) are held-by-default: the incumbent is the narrative owner and no challenger even attempted a distinct axis. The informative held cases are accounting and help desk, where a real challenger tried a same-axis pitch and it wasn’t enough.
How did the four pre-registered predictions score?
We wrote these down Tuesday, before coding a single axis. The honest scorecard:
| Pre-registered prediction | Result | Verdict |
|---|---|---|
| P1. Magnitude fails to separate crossed from held (no monotone size threshold) | Crossed spans 1.1x–160x; held accounting (2.2x) out-gaps crossed marketing automation (1.1x). No cut exists. | ✅ Confirmed (magnitude model refuted) |
| P2. Geometry reproduces the outcomes with ≤1 misclassification | 6 of 6 correct, 0 misses. Distinct-axis → crossed; same-axis → held. | ✅ Confirmed |
| P3. The magnitude/geometry discordant row is won by geometry | VPN — the study’s largest size gap (160x) — crossed; only the distinct privacy axis explains it. | ✅ Confirmed |
| P4. The held categories’ owners are same-axis | Xero (accounting) and Freshdesk (help desk) both coded same-axis — “better version” of the incumbent’s own axis. | ✅ Confirmed |
Four for four is a clean sweep — which is exactly why the next section, not the checkmarks, is the one to read.
The final ruling
Threading all four predictions back to the week’s question — when does an earned narrative beat an incumbent, and where does the moat hold?
Verdict: Conditionally yes — and the condition is geometric, not a matter of scale. The incumbency moat is a gate, not a wall: an earned narrative crosses it when it owns a distinct axis the incumbent can’t claim, and is held when it’s merely a better version of the incumbent’s own axis. Size doesn’t set the threshold.
- ❌ There is no size threshold — the biggest incumbent (160x) got crossed; a smaller one (2.2x) held. Out-sizing or being out-sized is not the deciding variable.
- ✅ Axis geometry decides it (6/6 on this set) — distinct axis crosses; same-axis “better version” is held.
- ✅ This completes last week’s finding — narrative ownership is earnable but not sufficient; now we know which earned narratives are sufficient: the ones on an axis the incumbent structurally can’t stand on.
- ⚠️ Held with a caveat — this is retrodiction on n=6 with only 2 crossed, so it’s a strong working rule, not a law (limits below).
The decision test: can your earned narrative cross the moat?
Turn the ruling into a check you can run before you spend a dollar on GEO. The whole question reduces to the geometry of the axis you’re competing on:
- Name your category’s defining attribute — the thing it’s bought for. “Privacy” for VPNs, “reporting” for accounting, “omnichannel” for help desks. This is the axis the engines’ pick tracks, because it’s who owns the defining language, not who has the most coverage (narrative ownership out-predicted raw mention volume 4-to-1).
- Ask the geometry question: is your narrative a distinct axis or a better version? The test is brutal and simple — could the incumbent credibly claim your attribute without abandoning its own identity? If yes, you’re on its axis (a “better version”) and the moat will hold. If the incumbent would have to become a different company to claim your attribute, you own a distinct axis and can cross.
- If you’re on a “better version,” stop and re-position. Cheaper, faster, more-omnichannel-than-the-omnichannel-leader — every one of these reinforces the incumbent’s frame. Out-mentioning it is wasted money: fewer than 1 in 5 brands earn both frequent mentions and consistent citations, and volume didn’t move the pick in any held category. Find a sharp sub-attribute you can own outright.
- If you own a distinct axis, earn it with dated proof, not a PR burst. Every crossed narrative traced to a concrete dated origin event — Mullvad’s founding-design privacy bet, its 2018 audit, its 2023 police raid that recovered no data; HubSpot’s free-CRM launch and app-marketplace flywheel. Ship a thing on a date, with a name, and get it into independent prose.
- Ignore the size gap entirely when you choose the fight. The data says a 160x size disadvantage is survivable on a distinct axis and a 2.2x disadvantage is fatal on a shared one. Your competitor’s review count is not the variable to optimize against.
The per-engine playbook: where to plant the proof
The geometry rule is platform-agnostic — own a distinct axis everywhere — but where you reinforce it shifts by engine, consistent with our earlier finding that there is no single universal GEO strategy across engines. Once you’ve picked a distinct axis, adapt the reinforcement:
| Engine | How to reinforce your distinct axis |
|---|---|
| ChatGPT (with search) | Leans on a small set of high-authority roundups and Wikipedia-grade consensus. Get your axis-defining event cited in the independent best-of lists it re-reads; a Wikipedia-referenceable proof (an audit, a public incident) carries disproportionate weight for anchoring your axis. |
| Perplexity | Most literal about attribute-tied prose and freshest sources — this is where owning the language of your axis pays fastest. Repeat the exact attribute phrasing across recent independent pages; recency and community mentions matter more here. |
| Google AI Mode / Gemini | Most incumbency-biased — the two categories where the incumbent held (accounting, help desk) look most like Google’s depth-and-authority preference. If you’re crossing a moat here, your distinct axis needs the deepest reinforcement: schema, a long independent trail, not one event. |
| Copilot (Bing) / Grok | Bing-indexed sources and, for Grok, real-time social. A dated axis-proof that gets discussed (Mullvad’s raid) travels well — make the event indexable and talked-about, not just published. |
What are the limits of this verdict?
We’d rather you distrust a 4-for-4 than take it on faith, so here’s what a clean sweep does and doesn’t earn:
- This is retrodiction, not prediction. The crossed/held labels were public before we coded the axes, so the coding could have been unconsciously fitted to the answer we knew. A clean 6/6 on labels you already know is weaker evidence than 6/6 on categories you’d never seen. This is the single biggest caveat, and it defines next week’s test.
- Axis coding is a judgment call. “Distinct” versus “better version” is a human read of the attribute and its origin evidence, not a number. A second independent coder could disagree on borderline rows; we coded blind to the size gaps to reduce the pull, but it’s still interpretation.
- n = 6, and only 2 crossed. Two crossed categories is a thin base to declare a law, and two of the four held cases are held-by-default with no distinct-axis attempt, so the informative held sample is really two.
- Single snapshot, correlational. One 2026 frozen roster, no engine re-query this cycle. Geometry classifies the outcomes; whether owning a distinct axis causes the cross — or both trace to the brand genuinely being different — is a causal question one observational week can’t close.
Next week’s question
The retrodiction caveat is the whole cliffhanger. A rule that scores 6/6 on categories whose outcomes you already knew has to prove itself on categories you’ve never scored — that’s the difference between a flattering description and a predictive law. So next week we run the honest test: take a fresh set of buyer-intent categories we’ve never touched, code each challenger’s axis geometry blind — before we know or look up which brand the engines pick — and register the prediction in advance. Then we query the engines and see whether distinct-axis challengers actually cross. If geometry predicts out-of-sample, the moat-gate model earns the word “law.” If it doesn’t, we’ll have caught ourselves hindsight-fitting, and we’ll say so. Same Lab discipline: hypothesis, method, data, verdict.
Frequently asked questions
Does a challenger have to be small to benefit from AI search?
No — and size isn’t the deciding factor at all. Our data shows the largest incumbent-vs-challenger size gap in the study (160x) got crossed while a much smaller 2.2x gap held. What decides whether a challenger beats an incumbent is the geometry of its narrative axis — whether it owns a distinct, orthogonal attribute the incumbent can’t claim — not how big or small either brand is.
What is a “distinct axis” versus a “better version”?
A distinct axis is an orthogonal trade-off the incumbent structurally can’t claim without abandoning its own identity — Mullvad’s privacy/no-logs against ExpressVPN’s speed-and-streaming. A “better version” is a comparative on the incumbent’s own axis — Xero’s “better reporting” when reporting is QuickBooks’ core territory. In our data, distinct-axis challengers crossed the moat 2 of 2 and “better version” challengers held 4 of 4.
How can Mullvad be the AI’s pick with 160x fewer reviews than ExpressVPN?
Because AI engines recommend based on how a brand is framed in independent editorial, not on raw review volume. Mullvad owns the privacy axis — nearly every independent mention ties it to audited no-logs and anonymity — so when an engine answers “most private VPN,” Mullvad is the reference point. ExpressVPN’s 28,215 reviews are about a different axis (speed, streaming). Axis ownership beat size.
If I own a distinct axis, am I guaranteed to win?
On this set, distinct-axis challengers crossed 2 of 2 — but we’re honest that it’s retrodiction on n=6, so it’s a strong working rule rather than a guarantee. It also has to be a real, earned axis backed by dated proof, not a slogan: every crossed narrative in our data traced to a concrete origin event that preceded the language spreading in independent prose. Next week we test the rule blind on fresh categories.
What should a challenger stop doing?
Stop trying to out-size the incumbent and stop selling a better version of its story. “Like the leader, but cheaper/faster/better” competes on the incumbent’s axis, and the engines keep picking the incumbent regardless of your review count. Spend instead on credibly owning a distinct axis the leader would have to abandon its identity to claim.
Where does this fit in the GEO Lab thread?
It’s the Friday verdict closing an arc that ran from Monday’s “does AI search level the field” pillar through Tuesday’s pre-registered method and Thursday’s results, and it builds on the prior finding that narrative ownership is earnable but doesn’t automatically buy the pick.
Sources
- GeoParrot GEO Lab, exp9 (this week): retrodiction/coding on the frozen 18-brand, 6-category roster — magnitude gaps (review-count proxies) and blind axis-geometry coding, scored against four pre-registered predictions; zero engines re-queried. First-party data.
- GeoParrot GEO Lab, prior weeks: narrative ownership out-predicts Share of Voice 4-to-1, the Mullvad reversal, and whether narrative ownership is earnable.
- Cure53, independent security audits of Mullvad VPN (2018, 2020, 2022–2025); public reporting on the 2023-04 Swedish police search of Mullvad’s office (no customer data seized).
- Vendor launch records: Zendesk Suite (2018-05), Freshdesk Omniroute (2019-02), HubSpot free CRM / App Marketplace (2014–2019).
- G2 and Trustpilot review counts (carried from the frozen roster); founding years via Wikipedia / company records.

Leave a Reply