The AI-Search Incumbency Moat Isn’t About the Incumbent’s Size — It’s the Challenger’s Narrative Axis. We Scored 6 Categories: Geometry 6-for-6, Magnitude Draws No Line

Quick answer: We scored all 6 frozen buyer-intent categories against the two rival moat models we pre-registered on Tuesday, and all four predictions held. Size didn’t draw the line: the biggest incumbent-vs-challenger size gap in the study — Mullvad’s 176 reviews against ExpressVPN’s 28,215, a 160x gap — got crossed, while a much smaller 2.2x gap in accounting held. There’s no size threshold that separates the two crossed categories from the four held ones. Axis geometry did draw it, 6 of 6: every challenger that owned a distinct axis the incumbent structurally can’t claim (VPN → privacy, marketing automation → free-CRM ecosystem) crossed the moat; every one selling a “better version” of the incumbent’s own axis (accounting, help desk, live chat, survey) stayed held. The decisive case is Mullvad — largest size gap in the set, yet crossed, and only geometry explains it. The honest catch: we already knew which categories crossed before we coded the axes, so this is retrodiction, not a fresh out-of-sample test — we flag the hindsight risk here and let this week’s verdict weigh it.

This is Thursday’s result in this week’s GEO Lab arc. Monday we asked whether AI search really levels the playing field for challengers and argued there’s still an incumbency moat that earned narrative crosses only sometimes; Tuesday we published the method and locked four predictions before scoring a single category — pitting a magnitude model (the moat is the incumbent’s size) against a geometry model (the moat is the shape of the axis you compete on). Today we run the numbers on the same frozen 18-brand, 6-category roster from the earlier weeks of this thread — zero engines re-queried, because re-querying would have broken the pre-registration.

Does the incumbent’s size explain who crosses the moat?

No — and not even close to a clean line. The magnitude model says a challenger crosses when the incumbent’s size advantage is small and stays held when it’s large. So we ranked each category by its size gap: the incumbency leader’s review count divided by the narrative owner’s. If size were the moat, the two crossed categories should cluster at small gaps and the held ones at large gaps. They don’t.

Lollipop chart on a log scale of each category's incumbent-vs-challenger review-count gap, coloured by outcome. VPN (crossed) 160.31x, accounting (held) 2.22x, help desk (held) 1.85x, marketing automation (crossed) 1.1x, survey (held) 1x no challenger, live chat (held) 1x no challenger. The crossed categories span 1.1x to 160x, and held accounting (2.2x) has a bigger gap than crossed marketing automation (1.1x), so no monotone size threshold separates crossed from held.

The crossed categories span the entire range — marketing automation at 1.1x and VPN at 160x — while the two held categories with a real challenge attempt sit right in between at 1.85x (help desk) and 2.22x (accounting). Held accounting has a bigger size gap than crossed marketing automation. There is no cut you can draw on this axis that puts the two crossed categories on one side and the held ones on the other. That’s our first pre-registered prediction — P1: magnitude fails to separateconfirmed, and the magnitude model refuted. Size is not the moat.

Does the challenger’s narrative axis explain it instead?

Yes — perfectly, on this set. The geometry model says the moat isn’t the incumbent’s size but the shape of the axis the challenger competes on. Coding each category’s owned attribute (blind to the size numbers) into two buckets: a distinct axis is an orthogonal trade-off the incumbent structurally can’t claim; a same axis is a “better version” of a dimension the incumbent already stakes. Geometry predicts distinct-axis challengers cross and same-axis challengers stay held.

Two-by-two confusion matrix of axis type versus outcome. Distinct axis and crossed: 2 categories (marketing automation, VPN). Distinct axis and held: 0. Same axis and crossed: 0. Same axis and held: 4 categories (help desk, live chat, survey, accounting). All 6 categories fall on the diagonal, zero misclassifications.

Every category lands on the diagonal. The two distinct-axis challengers — HubSpot (a free-CRM + app-marketplace ecosystem that email-automation leader ActiveCampaign doesn’t own) and Mullvad (audited no-logs privacy that speed-and-streaming leader ExpressVPN can’t claim) — both crossed. All four same-axis categories held. That’s 6 of 6 correct, zero misclassifications, clearing our pre-registered bar of at most one miss. P2: geometry reproduces the outcomes — confirmed. One honesty note: two of the four held categories (live chat, survey) are cases where the incumbent is the narrative owner — no challenger even attempted a distinct axis — so they’re held-by-default rather than held-against-an-attempt. The informative held cases are accounting and help desk, where a real challenger tried and a same-axis pitch wasn’t enough.

The decisive case: how did the study’s biggest size gap get crossed?

This is the test the whole week turned on. If magnitude and geometry ever disagree hard, one has to give — and VPN is the maximum-disagreement row. Mullvad carries 176 independent reviews against ExpressVPN’s 28,215: the largest size gap in the entire study, the one the magnitude model is most confident should hold. It crossed anyway. The engines’ pick is Mullvad.

Bar chart on a log scale of VPN review counts. ExpressVPN, the incumbency leader, has 28,215 reviews; Mullvad, the AI's pick, has 176 — a 160x magnitude gap. Annotation notes Mullvad owns the distinct axis of privacy and audited no-logs, which ExpressVPN's speed-and-streaming positioning cannot claim, so geometry rather than size decides the pick.

Magnitude says ExpressVPN by 160x. The engines pick the brand with 160x fewer reviews — because Mullvad owns a distinct axis: anonymous numbered accounts, cash payment, audited no-logs, a 2023 police seizure that recovered no customer data. That’s not a “better ExpressVPN.” It’s a different question — who protects your privacy? — and ExpressVPN’s speed-and-streaming positioning structurally can’t answer it. Only geometry explains this row, and it’s exactly the disagreement case the method was built to catch. P3: the discordant row is won by geometry — confirmed. This is the same reversal that has recurred all thread: a brand can be last on volume-style metrics yet first on what it owns.

Why do the “better version” challengers stay stuck?

Because a better version of the incumbent’s own axis reinforces the incumbent’s frame instead of escaping it. In accounting, Xero’s narrative is “better reporting and financial visibility” — but reporting is QuickBooks’ core territory, so “better reporting” is a comparative on QuickBooks’ axis, and the engines keep QuickBooks (a 2.2x size gap the challenger never overcame). In help desk, Freshdesk owns “omnichannel ticketing” — except Zendesk staked that same omnichannel territory earlier (Zendesk Suite in 2018 precedes Freshdesk’s Omniroute in 2019), so it’s the incumbent’s axis too. Both challengers were coded same-axis, and both held. P4: held owners are same-axis — confirmed. The pattern is consistent: to cross, you need an axis the incumbent can’t stand on, not a sharper version of the one it already owns.

How did the four pre-registered predictions score?

We wrote these down Tuesday, before coding a single axis. Here’s the honest scorecard.

Pre-registered prediction Result Verdict
P1. Magnitude fails to separate crossed from held (no monotone size threshold) Crossed spans 1.1x–160x; held accounting (2.2x) out-gaps crossed marketing automation (1.1x). No cut exists. ✅ Confirmed
(H0 magnitude refuted)
P2. Geometry reproduces the outcomes with ≤1 misclassification 6 of 6 correct, 0 misses. Distinct-axis → crossed; same-axis → held. ✅ Confirmed
P3. The magnitude/geometry discordant row is won by geometry VPN — the study’s largest size gap (160x) — crossed; only the distinct privacy axis explains it. ✅ Confirmed
P4. The held categories’ owners are same-axis Xero (accounting) and Freshdesk (help desk) both coded same-axis — “better version” of the incumbent’s own axis. ✅ Confirmed

Four for four is a clean sweep — which is exactly why the next section matters more than the checkmarks.

What are the limits of this result?

We’d rather you distrust a 4-for-4 than take it on faith, so here’s what a clean sweep does and doesn’t earn:

  • This is retrodiction, not prediction. The crossed/held labels were already public before we coded the axes, so the coding could have been unconsciously fitted to the answer we knew. We pre-registered this as limitation #1 on Tuesday and we’re not hiding from it: a clean 6/6 on labels you already know is weaker evidence than 6/6 on categories you’d never seen. The right test is a fresh set of categories coded blind — a candidate for a future week.
  • Axis coding is a judgment call. “Distinct” versus “better version” is a human read of the attribute and its origin evidence, not a number. A second independent coder could disagree on borderline rows; we coded blind to the size gaps to reduce the pull, but it’s still interpretation.
  • n = 6, and only 2 crossed. Two crossed categories is a thin base to declare a law. Two of the four held cases (live chat, survey) are held-by-default with no distinct-axis attempt, so the informative held sample is really two.
  • Single snapshot, correlational. This is one 2026 frozen roster with no engine re-query this cycle. Geometry classifies the outcomes; whether owning a distinct axis causes the cross — or both trace back to the brand genuinely being different — is the causal question Friday’s verdict has to rule on.

What does this mean if you’re a challenger brand?

The practical read is blunt: stop trying to out-size the incumbent, and stop selling a better version of its story. The categories where challengers crossed the AI-search moat weren’t the ones with small incumbents — the biggest incumbent in the study got crossed. They were the ones where the challenger owned a question the incumbent structurally can’t answer. If your positioning is “like the leader, but better/cheaper/faster,” you’re competing on the incumbent’s axis and the engines will keep picking the incumbent regardless of your review count. If you can credibly own a distinct axis — a trade-off the category leader would have to abandon its own identity to claim — that’s the shape that crosses. Whether that’s a durable, causal lever or a flattering description of brands that were already different is precisely what Friday’s verdict will decide, alongside a per-category playbook.

Frequently asked questions

What’s the difference between “magnitude” and “geometry” here?

Magnitude is the incumbent’s size advantage — how much bigger it is by reviews, list appearances, or ownership lead. Geometry is the shape of the axis the challenger competes on — whether it owns a distinct, orthogonal trade-off the incumbent can’t claim, or just a “better version” of the incumbent’s own axis. This week tested which one predicts whether a challenger becomes the AI’s pick, and geometry won 6-for-6 while magnitude drew no line.

How can Mullvad be the AI’s pick with only 176 reviews?

Because AI engines recommend based on how a brand is framed in independent editorial, not on raw review volume. Mullvad owns the privacy/no-logs axis of the VPN category — nearly every independent mention ties it to audited no-logs and anonymity — so when an engine answers “most private VPN,” Mullvad is the reference point. ExpressVPN’s 28,215 reviews are about speed and streaming, a different axis. Size lost to axis ownership.

Why didn’t you re-query ChatGPT, Perplexity, or Gemini this week?

Because we pre-registered the test on a frozen roster. The data source was locked on Tuesday; re-running the engines mid-experiment to look for a better answer is exactly the fishing pre-registration exists to prevent. This week’s “collection” was coding and magnitude computation on the already-frozen categories.

Does 4-for-4 mean the theory is proven?

No. It’s retrodiction on categories whose outcomes we already knew, with n = 6 and a judgment-based axis coding — strong enough to take the geometry model seriously, not strong enough to call it a law. The honest next step is testing it blind on categories we’ve never scored. We report the sweep and its limits in the same breath.

Where does this fit in the GEO Lab thread?

It’s the Thursday result in an arc that ran from Monday’s “does AI search level the field” pillar through Tuesday’s pre-registered method, and it builds on the earlier finding that narrative ownership is earnable but doesn’t automatically buy the pick. Friday’s verdict closes the week with the causal read and a per-category playbook.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *