Quick answer: After a week that ran from a claim to a locked bet to a blind score, here is the ruling in three parts. (1) The volume model is dead. Across the four categories where the biggest brand and the narrative owner are different companies, the volume leader took the best-overall crown 0 times out of 12 engine-observations. Being the biggest does not win the crown. (2) The narrative-crown model graduates — but we score it 4/4, not 5/6. The honest number is a perfect record on the four discriminating categories, because the two control categories could never separate the two theories in the first place; a control-arm miss is a footnote, not a fifth win to brag about. (3) Monday’s vanity-metric claim holds, but bounded. “Mention rate is vanity” is now backed by data — the volume leaders were named in nearly every answer yet crowned zero times — but that’s a six-category snapshot, so we’re retiring mention rate as a primary KPI, not declaring a universal law. The rest of this post is the reasoning, a per-engine checklist you can act on today, and next week’s test: can a crown be taken on purpose?
This is Friday’s verdict in this week’s GEO Lab arc. Monday we argued that “AI share of voice” is being sold as the new SEO rank, and it’s half a vanity metric — when the AI names almost everyone, being mentioned stops separating winners from losers, and the only score that moves a buyer is the best-overall crown. Tuesday we pre-registered six crown predictions and locked them before a single engine was queried. Thursday we scored them: 5 of 6, with the discriminating arm going a perfect 4/4. Today we rule on what that actually earns — and, per pre-registration discipline, we hold ourselves to the bands we set on Tuesday rather than the ones that would flatter us now.
What is the final verdict on the crown model?
Three separate claims went into this week. They get three separate verdicts — collapsing them into one number is exactly the mistake a share-of-voice dashboard makes.
| Claim under test | Verdict | What the data says |
|---|---|---|
| Volume model — the crown goes to the biggest brand (most installs, seats, sites, mentions) | ❌ Rejected | Volume leader crowned 0 of 12 engine-observations across the four disagreement categories |
| Narrative-crown model — the crown goes to the brand that owns the category’s defining sentence | ✅ Graduates (conditional) | 4/4 on the discriminating arm; narrative owner crowned 11 of 12. Conditional on n=6, single run, prose-coded crowns |
| Monday’s claim — mention rate is a vanity metric | ✅ Confirmed, bounded | Volume leaders appeared in nearly every answer (high “SOV”) yet won the pick 0 times. True at this scale; not proven universal |
The headline is not “5/6.” The headline is that two models entered and only one is standing — and the surviving one is the counterintuitive one, the one a smaller brand can actually act on.
Why score the model 4/4 instead of the flattering 5/6?
Thursday’s raw tally was 5 of 6 correct, which clears our pre-registered pass band (≥5 promotes, 4 is inconclusive, ≤3 forces a retraction). But a raw pass count is the wrong yardstick for a model, and here’s the crack we promised Thursday we’d settle.
The six categories were never equal. Four were disagreement categories — the volume leader and the narrative owner are different companies (Teams vs Slack, Adobe vs Figma, WooCommerce vs Shopify, WordPress vs Wix). Those are the only categories that can tell the two theories apart, because the two models predict different winners. The other two were controls — Zoom and Canva are both the volume leader and the narrative owner, so both models predict the same brand. A control can confirm the pipeline works, but it can never award a point to one theory over the other.
So the correct score for the narrative model’s core claim is its record where the claim is actually on the line: 4 out of 4. The fifth “hit” (Zoom) was a shared prediction that proves nothing about which model is right, and the one “miss” (Canva) was also a shared prediction — so it can’t dock the narrative model either. Reporting “5/6” would inflate the win with a non-discriminating point; reporting “4/6” would penalize the model for a category that couldn’t have discriminated anyway. The honest read is: on every category built to test it, the narrative-crown model went perfect, and the volume model went to zero. The controls are a footnote — a useful one, which is the next section.
Is the volume model really dead?
Yes, and this is the least hedged verdict of the three. Zero for twelve is not a close call. WordPress powers more of the web than any builder and was crowned by no engine; all three named Wix. WooCommerce runs on more stores than any hosted rival; all three named Shopify, with Google AI Mode calling it the “undisputed best overall e-commerce platform in 2026.” Microsoft Teams has far more seats than Slack through Office bundling; ChatGPT’s verdict was “🥇 Slack — best overall.” Adobe dwarfs Figma in revenue; Perplexity opened with “Figma is the best all-around UI design tool in 2026.”
The volume leaders didn’t disappear — they got demoted to a segment. Teams was “best if you already pay for Microsoft 365.” WooCommerce was “best for WordPress users who want maximum control.” That is the entire vanity-metric argument in one observation: a share-of-voice tool would have logged all four as present and prominent — high mention rate, healthy SOV — and completely missed that not one of them is the brand a buyer is pointed to. Mention rate counts the room; the crown counts the decision. When the two disagree this cleanly, optimizing for the room is optimizing the wrong thing.
Is Monday’s “mention rate is vanity” claim proven — or overstated?
It’s proven at the scale we tested, and we’ll say the boundary out loud rather than let the strong result run past its evidence. “Volume leader crowned 0 of 12” is a genuine, blind, out-of-sample result — it is not retrodiction, because the predictions were public before any query. That’s strong enough to change what you measure.
But it is six categories, one query per engine, one run. “Zero of twelve” could, on a different day, be “one or two of twelve” if a borderline answer flips — it would not resurrect the volume model, but it means the honest claim is directional and large, not 0.000 forever. So the verdict is calibrated: retire mention rate as a primary KPI — stop reporting “we’re mentioned in 80% of answers” as a win — but keep it as a diagnostic (if you’re not even mentioned, you’re not in the consideration set). The number that belongs on the dashboard is crown rate: in how many of your categories does the AI name you as the single best.
So what does the rewritten rule actually say?
Last week our geometry rule failed its own blind test at the appearance level — “did the challenger show up?” was too loose, because the segmented listicle lets almost everyone show up. This week we escalated the resolution from appearance to crown, and at that resolution the rule survives. Restated:
The AI’s best-overall crown goes to the brand that owns the category’s defining sentence, not the brand with the most volume — and it holds unless a fresh news event rewrites that sentence.
That last clause is earned by the one control that broke. Graphic design should have been a gimme (Canva is both biggest and default), but the crown split three ways because ChatGPT crowned Affinity outright, citing a live 2026 event: “the professional Affinity suite is now available free.” A crown is a strong prior, not a law — and a big enough event can override it. That’s not a weakness in the finding; it’s the finding’s most portable, most actionable edge, which sets up next week.
What’s the per-engine checklist to actually win a crown?
The three engines crown differently, and the raw answers show how. Here’s what to own for each — this refines our per-engine GEO checklist with this week’s crown data.
- ChatGPT (web-grounded) — win the news. It was the most event-sensitive engine: it flipped an entire category on a single “now free” event and used blunt single-winner language (“If I had to recommend just one: Affinity”). To take its crown you need a recent, citable reason to be the answer — a launch, a pricing change, a standout release the model can surface. Stale positioning loses to whoever made news last.
- Perplexity — own the quotable sentence. It opened categories with clean declaratives it lifted almost verbatim: “Wix is the best website builder for most people in 2026,” “Figma is the best all-around UI design tool.” Make sure authoritative sources state plainly “[You] is the best [category] for most people/teams in 2026” — Perplexity surfaces that exact line as the crown.
- Google AI Mode — own the consensus label. It crowned with status phrases (“undisputed best overall,” “종합 1위,” “절대적인 업계 표준”) and it hedged toward emerging tools where consensus was thin (UI design was our only 2/3). To hold Google’s crown you need the category’s consensus defining phrase repeated across many sources, not one loud claim — and to take it, exploit categories where the incumbent’s consensus is fraying.
30-second self-diagnosis: ask all three engines “best [your category] in 2026?” while logged out. If you get a “best for ___” slot but never the opening best-overall sentence, you are a volume/segment brand — and the fix is not more mentions. Pick the one sentence you want the AI to complete with your name, and make the web say it consistently.
What did we not settle this week?
- Durability. One run can’t separate a stable crown from a lucky snapshot. A re-run on a different day could wobble a 2/3 (like Figma) into a split. We don’t yet know how sticky these crowns are.
- Contestability. The Affinity control tells us a crown can move on a news shock — but not whether a brand can engineer that move on purpose, versus only benefiting from luck.
- Crown is a prose judgment. We coded each engine’s best-overall pick conservatively (majority-or-none), but a stricter or looser reading could move a borderline case.
- Small n. Four discriminating categories is a strong signal, not a census. The direction is clear; the exact magnitudes are not.
What’s next week’s test?
This week answered “who holds the crown?” Next week we escalate to the question every smaller brand actually cares about: can a crown be taken on purpose? The Affinity miss proved crowns move when a category’s defining sentence gets rewritten by an event. So we’ll pre-register a blind test around documented challenger moves — categories where a smaller brand has recently made a real narrative play — and predict, in advance, whether the crown flips. Pass, and narrative ownership becomes a lever you can pull, not just a trait you’re born with. Fail, and this week’s crowns are just incumbency we dressed up. We’ll also re-run two of this week’s categories to put a first number on durability. Same discipline: predictions locked before we query, bands set in advance.
Frequently asked questions
Why is the model’s score 4/4 and not the 5/6 you measured?
5/6 is the raw count of correct predictions, but two of the six categories were controls where the volume leader and the narrative owner are the same brand — so both competing theories predicted the same winner. Controls can confirm the test runs correctly, but they can’t award a point to one theory over the other. The narrative-crown model’s real score is its record on the four disagreement categories that actually discriminate between the theories, which is 4 out of 4. We report that number because it’s the honest one, not the flattering one.
Does the crown model mean a small brand can beat a big one in AI answers?
Yes — that’s the core finding. In every category where the biggest brand and the defining-sentence brand were different companies, the AI crowned the smaller narrative owner: Slack over Teams, Figma over Adobe, Shopify over WooCommerce, Wix over WordPress. Volume tells the model who is big; narrative tells it who is the answer. When you ask for the single best, it returns the answer, and narrative ownership is a lever a smaller brand can actually pull.
Should I stop tracking AI mention rate / share of voice entirely?
Demote it, don’t delete it. Mention rate is still a useful diagnostic — if you’re not mentioned at all, you’re not in the consideration set. But it’s a floor, not a goal, because our data shows the biggest brands are mentioned everywhere and crowned nowhere. Move crown rate to the top of the dashboard: in how many of your categories does the AI name you as the single best overall pick.
Why did graphic design miss, and does that weaken the verdict?
Graphic design was a control, not a discriminating test, so its miss doesn’t count against the narrative model’s core claim. It missed because a live 2026 event — the Affinity suite going free — split the crown three ways, with ChatGPT switching to Affinity outright. Rather than weakening the verdict, it sharpens it: a crown holds unless a fresh news event rewrites the category’s defining sentence. That’s the most actionable line of the week and the seed of next week’s test.
How confident should I be in a six-category result?
Confident in the direction, cautious about the magnitude. The predictions were locked publicly before any query, so this is a blind, out-of-sample result — not a rule fitted to its own data. That’s what makes “volume leader crowned 0 of 12” meaningful. But six categories and one run per engine is a snapshot, so treat it as a strong signal to change what you measure, not as a settled constant. Next week’s durability re-run will put a first number on how sticky these crowns are.
Sources
- GEO Lab experiment #11 — pre-registered predictions locked 2026-08-25, blind scoring 2026-08-27. Raw answer capture from ChatGPT (logged out, web-grounded), Perplexity, and Google AI Mode; neutral prompt “best [category] in 2026?”, one run each.
- This week’s arc: Monday — “AI share of voice” is a vanity metric, Tuesday — the pre-registered method, and Thursday — the blind results (5 of 6).
- Prior arc — the appearance-level test this one escalated: the 4-of-6 results and its verdict.
- Related GEO Lab work: scoring narrative ownership, why share of voice is the wrong metric, the incumbency-moat verdict, and the per-engine GEO checklist.

Leave a Reply