Quick answer: On Monday we found AI search is only half a level playing field — the classic incumbent moats (domain authority, a backlink hoard) dissolved, but a new one survived: being the category’s default narrative. Two of our six frozen categories saw a challenger cross that moat (a smaller, less-covered brand became the AI’s #1 pick); four held. This week we measure where the threshold sits — what separates a crossed category from a held one — by pre-registering two rival explanations on the same frozen 18-brand roster, with zero re-querying of any engine. (1) Magnitude — the moat is just how dominant the incumbent is: challengers cross weak incumbents (small lead) and lose to strong ones (big lead). This is the market’s implicit model. (2) Geometry — the moat is crossed only when the challenger owns a distinct attribute axis the incumbent can’t also claim (Mullvad’s “privacy,” HubSpot’s “inbound ecosystem”), and holds when the challenger merely offers a better version of the incumbent’s own axis (a cheaper omnichannel, a cleaner cloud report) — regardless of how small the incumbent’s lead is. The two disagree hardest on VPN, where the AI’s #1 pick has 176 reviews and beat a runner-up with ~28,000: magnitude says that crossing is impossible; geometry says a distinct axis is exactly where it happens. Four predictions are locked below, before we code a single attribute — including the four ways we’ve committed to declaring geometry wrong.
This is Tuesday’s design post, and it turns Monday’s slogan into a test: where is the incumbency-moat threshold — how strong does an incumbent’s default have to be before an earned narrative can’t cross it? On Monday we found the field is only half level: the door is open (about 85% of AI brand mentions are third-party, and web mentions out-predict backlinks ~3x), but the incumbent that is the category’s shorthand keeps a moat that survives the shift. The honest way to ask “where’s the threshold” is to first ask what the threshold is made of — the incumbent’s magnitude, or the geometry of the challenger’s narrative — and to commit to both, in writing, before we look.
What exactly are we measuring this week?
Two rival models of the same moat, stated plainly:
Hypothesis (H1 — geometry): The incumbency moat is crossed by axis distinctness, not incumbent size. A challenger crosses when it owns an attribute on a distinct axis the incumbent can’t also claim — the way Mullvad owns “privacy/no-logs/audited” and HubSpot authored “inbound” — and the moat holds when the challenger’s owned attribute is a graded better version of the incumbent’s own default axis (a cheaper omnichannel, a cleaner cloud report), even when the incumbent’s size lead is small. Coded across the roster, axis-distinctness should sort crossed from held; incumbent magnitude should not.
Null hypothesis (H0 — magnitude): The moat is simply the incumbent’s dominance. Crossings occur where the incumbent’s lead (size, list breadth, ownership share) is smallest; held categories are where it’s largest. If H0 holds, the challenger’s play is just “attack weak incumbents,” axis geometry adds nothing, and Friday’s verdict says the threshold is a size line.
The two worlds make opposite predictions about one category in particular. In VPN, the AI’s #1 pick is Mullvad — 176 reviews, on 1 of 8 best-of lists — beating a runner-up with roughly 28,000 reviews on 7 of 8. That’s the largest incumbent magnitude gap in the whole study. Magnitude predicts VPN is the last place a challenger could cross; geometry predicts it’s a clean crossing because Mullvad opened an axis (privacy) the incumbent never owned. The design below is built to force that disagreement into the open rather than hide it.
Why not just measure the incumbent’s size and call it the moat?
Because “how big is the incumbent” is the market’s model, and Monday’s data already dented it. The AI’s pick tracked who owns the defining language, not who has the most coverage — narrative ownership out-predicted raw mention volume 4-to-1. If the moat were pure magnitude, the tools that count mentions and share-of-voice would forecast crossings, and they don’t.
The confound we have to beat is that magnitude and geometry usually travel together: a giant incumbent tends to also own the category’s default axis, so in most categories you can’t tell which one is doing the work. To separate them you need the rows where they disagree — a big incumbent that got crossed anyway (magnitude says “held,” geometry says “crossed” if the challenger opened a new axis), or a small incumbent that held (magnitude says “crossed,” geometry says “held” if the challenger only offered a better version of the incumbent’s attribute). Those disagreement rows are where the threshold shows its true shape, and our frozen roster happens to contain them.
How do you separate magnitude from geometry without re-running an engine?
We reuse the exact frozen roster and corpus — 18 brands across 6 buyer-intent categories (help desk, live chat, marketing automation, survey, accounting, VPN), each already labeled pick, runner-up or challenger by two engines that independently agreed on the #1, from last month’s consensus study. The outcome — which brand each engine recommends — was frozen before we ever framed the moat question, so we can’t fish for it, and re-querying now would throw that away. What’s new this week is that we add two fresh codings to text we’ve already scored — an incumbent magnitude gap and an axis-distinctness label — and read them three ways:
| Test | The “magnitude” world (H0) predicts | The “geometry” world (H1) predicts |
|---|---|---|
| A · Crossed vs held The frozen outcome we’re explaining |
Read straight off the frozen labels — not a predictor. A crossing = the AI’s #1 pick is a challenger that is not the incumbency leader. | |
| B · Magnitude gap Is the moat just the incumbent’s size? |
Crossed categories have the smallest incumbent lead; held categories the largest. Rank by gap and the crossings sit at the bottom. | No relationship. A crossed category can carry the largest lead in the study (VPN). |
| C · Axis distinctness Is the moat about the challenger’s axis? |
Irrelevant — size decides the moat, not the shape of the narrative. | Every crossed category is distinct-axis; every held category is same-axis (or the incumbent is also the owner). |
The tests do different jobs. A fixes the outcome we owe an explanation for. B gives the market’s model its fair, pre-committed shot. C states our rival predictor with a rule locked in advance — and the whole point is to see which one retrodicts the frozen labels with fewer errors.
Test A — Which categories crossed the moat? (the frozen outcome)
A category crossed when the AI’s #1 pick is a challenger that is not the incumbency leader — a smaller or less-covered brand won the recommendation over a bigger one. It held when the pick is the incumbent default, or when the incumbent is also the narrative owner (nobody crossed anything). These labels come straight off the frozen exp7 roster, with no re-query. We’re transparent that the outcome is already known — Monday’s post named the crossings — so this week is retrodiction, not forecasting: the test is whether a pre-committed rule (magnitude, or geometry) matches labels we’re not allowed to change. That’s a real risk of hindsight, which is exactly why every coding rule and decision threshold is fixed now, before we score one attribute, and why the person coding axis-distinctness does it from the corpus language while blind to the size gaps.
Test B — Is the moat just the incumbent’s size? (the magnitude gap)
For each category we compute the incumbency leader’s dominance gap over the crossing (or narrative-owning) challenger on the frozen proxies: the review-count ratio (market size), the list-breadth difference (how many of the best-of lists each brand appears on), and the ownership-share lead carried over from last month. Then we rank all six categories by that gap. Under the magnitude model, the crossings must cluster at the small-gap end — a challenger can only topple an incumbent it’s close to. The stress case is VPN: the crossing brand (Mullvad, 176 reviews, 1 of 8 lists) sits behind a runner-up with ~28,000 reviews on 7 of 8 — the biggest gap on the board. We pre-register the exact test: if the two crossed categories are not among the three smallest magnitude gaps, magnitude fails as the threshold.
Test C — Is the moat about the challenger’s axis? (axis-distinctness coding)
For each category we name the incumbent’s default attribute (the language independent writers use as the category’s shorthand) and the challenger’s owned attribute (from the frozen ownership scores), then code the relationship as same-axis or distinct-axis using a rule fixed in advance:
- Same-axis — the challenger’s attribute is described in the corpus as a comparative of the incumbent’s: cheaper, simpler, more of the same capability. Freshdesk as “more affordable omnichannel” (omnichannel being the help-desk default Zendesk already anchors); Xero as “cleaner cloud reporting” (reporting being the accounting default). The challenger is grading the incumbent’s own axis.
- Distinct-axis — the challenger’s attribute is an orthogonal selection criterion that trades off against the default rather than grading it. Mullvad’s “privacy / no-logs / audited” is not a better version of a mainstream VPN’s “speed and streaming unblocking”; it’s a different reason to choose, and the corpus treats them as a trade-off (“fast but less private,” “private but fewer servers”) rather than a scale.
We code every category same- or distinct-axis, log the corpus sentences that decide each call so a reader can check the classification themselves, and flag any marginal judgment out loud. The pre-registered classifier is simple: crossed should equal distinct-axis, held should equal same-axis.
What did we lock before coding a single attribute?
Pre-registration means the predictions are on the record now, before we compute one gap or code one attribute, so we can’t quietly reshape them on Thursday:
- Magnitude fails as the threshold. At least one crossed category carries a larger incumbent magnitude gap than at least one held category — so no size line cleanly separates the two groups. If crossed and held instead sort neatly by gap, magnitude wins and geometry is unnecessary; we’ll say so.
- Geometry classifies. Coding each category same- vs distinct-axis retrodicts crossed/held with at most 1 error across the 6: every crossing is distinct-axis, every hold is same-axis (or incumbent-owned). Two or more errors falsifies geometry as the threshold.
- Geometry wins the disagreement rows. In every category where magnitude and geometry predict different outcomes, the frozen label matches geometry. VPN is the decisive one: largest gap yet crossed — geometry must have coded it distinct-axis and magnitude must have called it held.
- The held-owner tell. In held categories where a challenger owns the narrative but still loses (accounting: Xero owns reporting, QuickBooks wins; help desk: Freshdesk owns omnichannel, Zendesk wins), the owned attribute must code as same-axis — the challenger earned the incumbent’s own attribute, not a new one. If a losing owner’s attribute codes distinct-axis, geometry is wrong and Friday’s verdict says the threshold isn’t about the axis after all.
If the data contradicts these — crossings that sort cleanly by size, axis coding that misses the labels, a held owner sitting on a genuinely distinct axis — we publish the contradiction and Friday tells challengers the moat is about magnitude, not narrative shape. A method that can only confirm its own hypothesis isn’t a method; it’s a sales page.
Where could this break? (the honest limitations)
- We’re retrodicting a known outcome. The crossed/held labels were frozen last month and are already public, so this is explanation, not forecasting — which invites hindsight bias. We mitigate by fixing every coding rule and decision threshold now, coding axis-distinctness blind to the magnitude gaps, and logging the deciding corpus sentences; but a reader should weigh this as a natural-experiment retrodiction, not a prediction from zero.
- Axis-distinctness is a judgment call. “Same axis vs distinct axis” is coded by reading corpus language, and reasonable people can disagree on the marginal cases — is HubSpot’s “ecosystem/inbound” a genuinely new axis, or just better integration? We pre-commit the rule (a comparative-of the incumbent’s attribute vs a criterion that trades off against it), publish the sentences behind each call, and flag every marginal one rather than burying it.
- n = 6, with only 2 crossings. Two crossed categories can’t carry a strong statistical claim, and a single miscoded category can flip the result. We report this as a pattern with a mechanism, not a law, and the small-sample caveats from the reused roster ride along.
- Magnitude has many faces. Size could be reviews, list breadth, revenue, headcount, or age; we pre-register reviews + breadth + ownership-lead, but a different size proxy might separate crossed from held where ours doesn’t. If any reasonable magnitude proxy draws a clean line, we’ll report it and soften the geometry claim accordingly.
- It’s a single frozen snapshot. One point in time, two engines. A category that reads “held” today could cross next quarter if a challenger opens an axis, and vice versa. The threshold we’re mapping is where it sits now — a description of the current terrain, not a constant of nature.
FAQ
What is the “incumbency moat” in AI search, and what’s its threshold?
The incumbency moat is the advantage a category leader keeps after domain authority and backlinks stop mattering: being the brand independent writers use as the category’s shorthand, so the model reaches for it by default. The threshold is the question of how strong that default has to be before an earned challenger narrative can’t cross it. This week we test whether the threshold is set by the incumbent’s size (magnitude) or by whether the challenger owns a distinct attribute axis (geometry).
Is beating a big incumbent in AI search just about picking a weak category?
That’s the magnitude model, and it’s exactly what we’re testing against. Our decisive counter-case is VPN, where the AI’s #1 pick (176 reviews, 1 of 8 lists) beat a rival with ~28,000 reviews on 7 of 8 — the largest size gap in the study. If magnitude decided the moat, that crossing shouldn’t exist. We pre-register that geometry — owning a distinct axis like “privacy” — explains it where size can’t.
What is a “narrative axis,” and how is it different from a feature?
An axis is the dimension of choice a category is decided on — speed, privacy, ease of use, price. A feature is a spec; an axis is the reason a buyer (and a model) picks one brand as the answer. A challenger owns a distinct axis when its attribute trades off against the incumbent’s rather than grading it: “private but fewer servers” is a distinct axis from “fast and unblocks streaming,” whereas “cheaper omnichannel” is the same axis as the incumbent’s omnichannel, just graded better.
Why reuse last month’s roster instead of running a fresh study?
Because the outcome — which brand each engine recommends — was frozen before we framed the moat question, so we can’t retrofit it, and re-querying the engines now would destroy that. We’re adding two new codings (a magnitude gap and an axis label) to text we’ve already scored, which keeps every result comparable to the correlation we’re trying to explain. The cost is a small six-category sample; we treat the findings as strong signals, not laws.
When do the results come out?
We code the magnitude gaps and axis labels midweek and publish the full data — the ranking, the classifications, the corpus sentences behind each call — on Thursday, then rule on Friday: is the incumbency-moat threshold about the incumbent’s size or the challenger’s axis, and what does that mean per category? The four predictions above are on the record now so you can hold us to them, including the ways we’ve committed to declaring geometry wrong.
Sources
- GEO Lab — “AI Search Levels the Playing Field.” Half True — There’s an Incumbency Moat (Monday’s fact-check: the door is open but the default-narrative moat survives, and earned narrative crosses it only sometimes)
- GEO Lab — Is Share of Voice the Wrong AI-Visibility Metric? The Verdict (the pick tracks attribute-tied prose, not mention volume — narrative ownership out-predicts share of voice 4-to-1)
- GEO Lab — Is “Share of Voice” the Wrong Way to Measure AI Visibility? (the Mullvad case: the least-incumbent brand owning its category’s attribute and winning the pick)
- GEO Lab — We Scored the AI Consensus Pick Against Its Runner-Ups (the frozen 18-brand roster and the incumbency data — reviews, list breadth, ownership — we reuse)
- GEO Lab — Can a Challenger Deliberately Earn Narrative Ownership? Our Pre-Registered Causal Test (last week’s design and the attribute-ownership scoring this test builds on)
- GEO Lab — Does Narrative Ownership Out-Predict Share of Voice? (the 4-to-1 result and the ownership-share measure carried into this test)
- GEO Lab — How Does a Brand Become the AI “Consensus” Pick? (why consensus is real across engines and can’t be self-manufactured)
- GEO Lab — No Universal GEO Strategy — The Per-Engine Checklist (the playbook format Friday’s verdict will follow)

Leave a Reply