Quick answer: This is a pre-registration, not a result. On Monday we argued that AI “share of voice” is drifting into a vanity metric — 2026 “best [category]” answers are segmented listicles (“best for beginners,” “best free,” “best for teams”) that name nearly every credible tool, so mention rate maxes out and stops sorting winners. The one score that still sorts is the best-overall crown: who the engine names first, “best overall,” or “if you only pick one.” Today we test whether that crown is predictable — blind, on categories we’ve never queried. Two rival models compete: the volume model (crown follows the biggest install base / most mentions) versus our narrative model (crown follows whoever owns the category’s defining sentence, regardless of raw size). We locked 6 predictions across 6 fresh categories — and in 4 of them the volume leader is deliberately not the narrative owner (Microsoft Teams vs Slack, Adobe vs Figma, WooCommerce vs Shopify, WordPress vs Wix). In each of those we bet the crown goes to the narrative owner. If it does, mention/volume is confirmed as a vanity signal and the narrative model graduates; if the crown keeps going to the volume leader, our thesis is wrong and we’ll say so. Wednesday we query blind, Thursday we score, Friday we rule. Pass at 5 of 6; retract at 3 of 6 or worse.
This is Tuesday in this week’s GEO Lab arc, and it’s built to be uncomfortable — the same discipline we used last week. Monday’s fact-check made a claim: the crown, not mention rate, is the decision-relevant metric. A claim that only describes data you already have is cheap. The honest follow-through is to make it predict something it has never seen, with the answer written down first. That’s a pre-registration — a public, timestamped commitment that removes our freedom to reinterpret the outcome after it lands.
What exactly are we predicting — and why the crown, not mention rate?
Last week we ran a blind test on the wrong yardstick and learned it the hard way. We predicted whether six challengers would appear in the engines’ answers — and all six appeared in all three engines, 18 of 18. On the appearance yardstick a coin flip would have scored the same, because in 2026 almost everyone appears. The wreckage pointed straight at a scarcer target. This week we move up a level and predict the crown: the single best-overall pick in a neutral “best [category]” answer. Because the crown is scarce, a prediction about it can actually be wrong — which is the only kind of prediction worth making.
What is the “narrative model,” stated as a mechanical prediction?
All summer we’ve circled one idea — that AI recommendations track who owns a category’s defining narrative, not who is biggest. Here we compress it into a rule that produces one binary call per category, with no wiggle room:
- Narrative model (our hypothesis). The crown goes to the brand that owns the category’s defining sentence — the one attribute a neutral buyer treats as the category’s core purpose (“the app that is team chat,” “the tool you start a store with”). Size, install base, and mention count are not inputs.
- Volume model (the null we’re testing against). The crown goes to the brand with the largest footprint — most installs, most reviews, highest mention rate. This is the worldview baked into the “share of voice” dashboards we questioned on Monday.
When these two models point at the same brand, the category tells us nothing — either model “wins.” The test only has teeth where they disagree: categories where the volume leader is clearly not the narrative owner. So we built the roster around disagreement on purpose.
How did we pick the categories — and what makes this “blind”?
We fixed a selection rule before choosing anything, so we couldn’t cherry-pick easy wins:
- Fresh. Six categories, none of which appeared in our last two frozen rosters (help desk, accounting, live chat, marketing automation, survey, VPN — and last week’s browser, analytics, note-taking, password manager, email, CRM). These are cases neither model has touched.
- Disagreement by design. Four of the six are categories where the volume leader is demonstrably not the narrative owner — the only cases that can separate the two models. The other two are agreement controls where volume and narrative point at the same brand, included to confirm the test harness itself works.
- Clear structure. Each category has a recognizable volume leader and a recognizable narrative owner that GEO practitioners would name without hesitation.
“Blind” means the codes below were written from general category knowledge — what each brand is known for and roughly how big it is — not from any AI output. We have not queried ChatGPT, Perplexity, or Google AI Mode for “best [category]” on any of these six. We know these brands’ reputations; what we do not know is which one each engine will actually crown this week. That gap is exactly what the test measures.
The six pre-registered crown predictions (locked)
This table is the frozen artifact. Each row was coded before collection; the rationale is logged now so it can’t be rewritten later. In every row our prediction is the narrative owner.
| Category | Volume leader | Narrative owner | Why they diverge | Type | Predicted crown |
|---|---|---|---|---|---|
| Team chat | Microsoft Teams | Slack | Teams wins on Office bundling (seats); Slack defined “team chat” as a category (distribution ≠ definition) | Disagreement | Slack |
| Product / UI design | Adobe | Figma | Adobe wins on total creative-suite install base; Figma owns “design in the browser, together” (suite size ≠ category default) | Disagreement | Figma |
| E-commerce platform | WooCommerce | Shopify | WooCommerce has more installs (free WP plugin); Shopify owns “start and run an online store” (install count ≠ the buying decision) | Disagreement | Shopify |
| Website builder | WordPress | Wix | WordPress powers ~43% of the web (volume) but reads as a technical CMS; Wix owns “build your own site, no code” (market share ≠ the builder narrative) | Disagreement | Wix |
| Video conferencing | Zoom | Zoom | Volume and narrative agree — “hop on a Zoom” is the category (control: harness check) | Agreement | Zoom |
| Graphic design (non-designer) | Canva | Canva | Volume and narrative agree — Canva owns both usage and “design anything, no skills” (control: harness check) | Agreement | Canva |
Four disagreement cases, two agreement controls. The four disagreement rows carry the test: if the narrative model is real, the crown in each should go to Slack, Figma, Shopify, and Wix — not to the bigger Teams, Adobe, WooCommerce, or WordPress. The two controls should crown Zoom and Canva; if they don’t, something is wrong with the collection, not the hypothesis. Where a category’s coding was genuinely arguable (e.g. website builders where Squarespace could rival Wix as the narrative owner), we still committed to a single brand rather than hedging — the judgment is now public and unchangeable.
How will Wednesday’s blind collection work?
To keep the outcome uncontaminated, the protocol is fixed here, in advance:
- One neutral query per category. “What is the best [category] in 2026?” — asked verbatim, logged in full. The neutral question is the fair, hard bar: the narrative owner has to win the crown without us framing the question around its axis.
- Three engines. ChatGPT (web search on), Perplexity, and Google AI Mode — the same three we’ve used since our cross-engine consensus work, logged out, single run.
- “Crown” defined now. The crowned brand is the one the engine names as best overall, first, or “if you only pick one.” A “best for beginners”/”best free” slot is not the crown, and a passing mention does not count. If an engine gives a segmented list with no clear top pick, that engine’s crown is recorded as “none.”
- Scoring rule. A category’s observed crown is the brand that takes best-overall in at least 2 of 3 engines. A prediction is correct when the observed crown equals the predicted brand. If the three engines split with no 2-of-3 majority, the observed crown is “none” and that prediction is scored as a miss (the conservative call).
- No touching the codes. Wednesday’s collector records raw engine output and the best-overall pick only; it does not see or edit this table.
What result would prove the narrative model wrong?
The pass/fail line is committed now, so Thursday’s write-up can’t move it:
- 5 or 6 of 6 correct → the narrative model graduates from a Monday argument to a working (still provisional) out-of-sample crown predictor — and mention/volume is confirmed as a vanity signal for who actually wins.
- 4 of 6 → inconclusive. Suggestive, not confirmed — reported as not passing.
- 3 of 6 or worse → we retract. Three correct is what a coin flip yields on six calls; at that level we say plainly that the volume model predicts the crown just as well, and Monday’s “mention rate is vanity” claim was overstated — in public, in the same place we made it.
Because the two controls are near-certain, the real signal lives in the four disagreement cases, so we’ll also report that subset on its own: the narrative model only truly earns the word “model” if the crown goes to the smaller narrative owner in at least 3 of those 4. Concretely, the model loses if Microsoft Teams is crowned “best team chat” over Slack, if Adobe outranks Figma for design, if WooCommerce beats Shopify, or if WordPress is named the best website builder over Wix. We are naming those failure modes now, before we can see which way they break.
What are the limits of even a passing result?
A clean pass would be real evidence, but we won’t oversell it:
- Small and binary. n=6, one crown each, with four load-bearing cases. Encouraging, not decisive.
- Selection judgment. We chose the categories and which brand is the “narrative owner.” A skeptic can argue we picked divergences we already knew would break our way; the locked codes and named failure modes are our guardrails, not a full defense.
- Coding judgment. “Volume leader vs narrative owner” is a human call, and a few categories (website builder) have a plausible second narrative owner. Committing to one brand publicly, before results, is the only real check.
- Snapshot and personalization. Engines change weekly and personalize; this is a single logged-out run. A pass says the model predicted this snapshot, not that it holds forever — the no-universal-strategy caveat still applies.
- Prediction is not mechanism. Even 6-for-6, this shows narrative forecasts the crown — not that owning the defining sentence causes the engine to crown you. The causal claim stays out of scope, as it did in our earnability verdict.
That’s the honest frame for the week. We took the standard we applied to every other GEO finding and pointed it at Monday’s own claim. If narrative calls these six crowns blind, it earns the word “model.” If it doesn’t, you’ll read that here on Friday — the same place we made the claim. You can see our full first-party scorecard in the 2026 GEO Benchmark.
Frequently asked questions
What’s the difference between this week’s test and last week’s?
Last week we predicted whether a challenger would appear in the answer — and everyone appeared, so the yardstick couldn’t fail anyone. This week we predict the crown: the single best-overall pick. The crown is scarce, so a prediction about it can actually be wrong, which makes it a real test.
Why pit “narrative” against “volume” specifically?
Because the volume model is the assumption behind the “AI share of voice” dashboards we questioned on Monday — bigger footprint, higher mention rate, therefore winning. If the crown instead follows the narrative owner even when they’re smaller, then mention/volume is a vanity signal for who actually gets recommended. The four disagreement categories are where those two worldviews give different answers.
Isn’t picking the categories yourselves a way to rig the result?
It’s the biggest risk, which is why we fixed a selection rule (fresh categories, a clear volume-vs-narrative split) and locked a single predicted brand per row with the failure modes named. An independent category list would be stronger; the public, timestamped codes are the guardrail against reinterpreting a near-miss later.
Why the neutral “best [category]” query instead of an axis-framed one?
Because “best team chat app” or “best no-code website builder” stacks the deck for the narrative owner. The neutral “best [category] in 2026?” is the harder, fairer test — the exact question a real buyer asks — and the narrative owner has to win the crown on it unaided.
How can I run a version of this on my own category?
Before you query anything, write down two names: the volume leader in your category and the brand that owns its defining sentence. Predict the crown goes to the narrative owner, then run “best [category] in 2026?” on ChatGPT, Perplexity, and Google AI Mode and check who’s named best overall. One honest out-of-sample call beats a dashboard full of mention-rate percentages — which was the whole point of Monday’s piece.
Sources
- GeoParrot GEO Lab — “AI Share of Voice” is a vanity metric; the crown is the score that moves a buyer (this week’s Monday fact-check).
- GeoParrot GEO Lab — last week’s pre-registration and blind results (all six challengers appeared) (the appearance yardstick that couldn’t fail).
- GeoParrot GEO Lab first-party data — who owns a category’s defining narrative, earnability verdict, the incumbency moat is a geometry gate, cross-engine consensus results.

Leave a Reply