Quick answer: Conditionally yes — Share of Voice is the wrong ranking signal, and you should replace it, but not with what its critics call “narrative.” Across 6 buyer-intent categories on our frozen first-party roster, the metric 2026 GEO tools optimize — best-of-list breadth (your share of the lists that name you) — matched the AI’s actual pick in only 1 of 6 categories. A scored narrative-ownership metric — a brand’s share of its category’s defining-attribute prose — matched the pick in 4 of 6, four times better, with a clean tier ladder (picks own 44% of the attribute-language, runner-ups 35%, challengers 22%) and a textbook reversal in Mullvad (last on breadth, first on narrative, and the engines’ #1 VPN). But our self-falsification test failed: we bet framing (how centrally a brand is described) would out-predict frequency (how much attribute-tied prose exists) — and frequency won, 5 of 6 to 3 of 6. So the honest verdict is: stop optimizing list breadth; start earning the volume of independent prose that ties you to the one thing your category is bought for. Below: the full ruling, a per-engine checklist, and next week’s question — can a challenger earn that on purpose, or only inherit it?
This is Friday’s verdict, closing this week’s GEO Lab arc. Monday we showed Share of Voice ranks the AI’s runner-up ahead of the winner and named “narrative ownership” as the lever. Tuesday we pre-registered a method and four predictions before scoring a single mention. Thursday we ran the numbers: three of four predictions held, one failed. Today we rule on the week’s question, turn the result into a checklist you can act on, and hand off the thread to next week.
What exactly were we ruling on this week?
One question, in two halves, threaded through four posts:
- (A) Is Share of Voice the wrong way to measure AI visibility? The 2026 GEO tooling wave — Profound, AthenaHQ, Peec and a dozen others — is built on it: your share of AI brand mentions, reported as a “mention rate” you’re told to grow.
- (B) Should you quantify “narrative ownership” instead? Monday’s fallback lever — owning the language of your category’s defining attribute in independent prose — was a slogan until we scored it this week.
A verdict has to answer both without rounding in our own favour. Here it is, split the way the data splits.
Verdict on (A): Is Share of Voice wrong? ✅ As a ranking signal, yes.
Confirmed, with one precision. Share of Voice fails as a ranking dial and survives only as a presence floor:
| What Share of Voice is used for | Verdict | The evidence |
|---|---|---|
| Ranking dial — “grow your mention share to win the recommendation” | ❌ Wrong | Best-of-list breadth (the count SoV tools optimize) named the AI’s actual pick in just 1 of 6 categories; runner-ups sat on as many or more lists than the picks they lost to |
| Presence floor — “am I in the conversation at all?” | ✅ Valid | A brand generally needs some independent presence to be considered — a pass/fail check, not a dial you optimize toward |
The distinction is the whole point. If you spend next quarter closing your Share-of-Voice gap — chasing more list placements to match a competitor’s breadth — our data says you may be copying the brand that lost the recommendation. Best-of breadth is the number that’s easy to sell precisely because it’s easy to count; it just doesn’t track the outcome you care about.
Verdict on (B): Is “narrative ownership” the replacement? ⚠️ Yes — but rename it, because framing isn’t the engine.
A qualified win, and the qualification matters more than the win. Scored narrative ownership out-predicted Share of Voice decisively — but why it works is not what we bet:
| Pre-registered claim | Result | What it means for the metric |
|---|---|---|
| Picks out-own their runner-ups on attribute-language (≥4 of 6) | ✅ Held — exactly 4 of 6 | The metric separates tiers where breadth couldn’t: 44% / 35% / 22% for pick / runner-up / challenger |
| A Mullvad-type reversal exists (breadth-last, narrative-first, and the pick) | ✅ Held | Mullvad: 1 of 8 VPN lists yet 43% of “privacy” prose — first, and the engines’ #1 |
| Narrative ownership out-predicts Share of Voice | ✅ Held — 4 of 6 vs 1 of 6 | The headline: the metric worth watching is a share, but of the right text |
| Framing beats frequency (centrality > raw count) | ❌ Refuted — raw count 5 of 6, framing 3 of 6 | The predictive engine is largely volume of attribute-tied prose, not how cleverly you’re framed |
So “narrative ownership” is real and it beats Share of Voice — but calling it “narrative” oversells the framing half. The signal that actually predicts is Attribute-Tied Prose: how much independent editorial talks about you in the same breath as your category’s defining attribute. It is still a count — which is why we won’t pretend it’s the opposite of Share of Voice. The difference from Share of Voice isn’t “prose vs counting.” It’s what you count: SoV counts list placements (1 of 6); this counts attribute-tied prose (4–5 of 6). Same act, right target.
Framing still earns its keep in one place: the underdog reversals. Raw count predicts everything except VPN, and VPN breaks toward Mullvad only because 82% of its mentions tie it to “no-logs / RAM-only / anonymous” — the volume-independent centrality that lets a low-volume brand win. So the layered ruling: volume of attribute-tied prose is the primary signal; centrality is the tiebreaker that decides the low-breadth upsets.
So what’s the one metric to replace Share of Voice? (Attribute-Tied Prose Share.)
If you track a single AI-visibility number in 2026, make it this, not mention breadth:
Attribute-Tied Prose Share (ATPS): of all the independent-editorial sentences that describe your category’s defining attribute — the one thing “best [category]” is really bought for — the share that name your brand. Freeze the attribute from a neutral “how to choose” guide (never from yourself), then measure your share of the prose that carries it. In our data this ranked the AI’s pick #1 in 4 of 6 categories; best-of-list breadth managed 1 of 6.
Two things the failed prediction forces us to add, so you don’t misread the metric:
- Volume is doing most of the work. You can’t win ATPS with a handful of perfectly-worded placements. The challenger’s job is to earn more independent writing that ties them to the defining attribute — not a few flawless mentions. Framing alone is the exception (Mullvad), not the plan.
- Count the right prose, not any prose. Generic mention volume is closer to Share of Voice. What predicted the pick is mentions welded to the attribute — “SurveyMonkey, for reaching respondents anywhere,” not “SurveyMonkey is popular.”
What’s the per-engine checklist? (How to earn attribute-tied prose where it lands.)
One reason a single metric works across engines: we’ve shown the engines cite almost entirely different source URLs (about 0.5% overlap) yet converge on the same recommendation. The only thing that survives across all those different pages is how the brand is described. That dictates the playbook’s first rule — and the rest is per-engine.
| Engine | What it reads | Your move for attribute-tied prose |
|---|---|---|
| Rule 0 — all engines | Different URLs, same conclusion | You can’t buy this on one page. Seed the same attribute-shorthand consistently across community, editorial and reviews — the language is the only thing that travels between engines. Volume across sources, not one perfect placement. |
| Perplexity | Live web search, Reddit-heavy citation mix, strong freshness bias | Earn attribute-tied prose in community threads and fresh editorial. The defining-attribute phrase should appear in real Reddit/forum answers, not just your own pages — and recency counts, so keep it flowing, don’t set-and-forget. |
| ChatGPT (search on) | Parametric memory blended with browsing; leans on long-standing, widely-syndicated editorial | Get the attribute association into evergreen, high-distribution editorial and keep the phrasing consistent over time. ChatGPT rewards the brand that has been the reference point, not a fresh burst. |
| Google AI Overviews / Gemini | Google’s index + review corpus — but breadth alone didn’t predict | Not “be on more lists.” Be framed for the attribute in the editorial Google already trusts, and in review prose that names the attribute (not just star counts — rating didn’t predict the pick either). |
| Microsoft Copilot / Bing | Bing’s index — a different URL draw from the others | Don’t assume a ChatGPT win carries over. Because the source overlap is tiny, deliberately place attribute-tied prose in Bing-indexed sources too; treat each engine’s corpus as its own coverage gap to close. |
Notice what’s absent from every row: “get on more best-of listicles.” That’s the breadth play that scored 1 of 6. The consistent instruction is to make independent writers reach for you when they explain the attribute — across as many of the pages each engine happens to read as you can reach.
What are the honest limits on this verdict?
We’re ruling on a small, hand-coded study, and we won’t dress it up:
- Correlational, not causal — this is the big one. Everything above shows attribute-tied prose predicts the pick. It does not prove a challenger can manufacture it and move the engines. Ownership may be a byproduct of genuinely being the best tool. That gap is exactly next week’s question.
- Coders disagreed ~36% of the time (67% on survey). Read the per-category numbers as directional, not decimal-precise.
- Six categories, 18 brands, one frozen roster. Prediction 1 landed exactly on its 4-of-6 line — thin margins. We’d want a fresh category set before calling ATPS a law rather than a strong signal.
- The attribute freeze is a discipline, not a proof. We froze each attribute from a neutral guide and published the sources so you can object and re-run — but retrofitting remains the failure mode we can’t fully rule out from the outside.
What’s next week’s question?
This week answered prediction: attribute-tied prose predicts the AI’s pick better than Share of Voice. The question it leaves open — flagged in every post this week — is action:
Next week’s GEO Lab thread: Can a challenger deliberately earn narrative ownership — or is it locked in by category history? If we take a brand that doesn’t own its attribute-language today and seed attribute-tied prose across the sources each engine reads, does its ATPS — and eventually the engines’ pick — actually move? Or is owning “the private one,” “the all-in-one,” “the easy one” a slow inheritance you can’t buy your way into? That’s the difference between a metric you watch and a lever you pull — and it’s what we’ll try to test next.
Until then, the practical instruction from this week stands: drop best-of breadth as your visibility KPI, and start measuring — and earning — the share of your category’s defining-attribute prose that names you.
Frequently asked questions
So is Share of Voice a bad metric or not?
It’s the wrong ranking signal and a fine presence check. In our data, best-of-list breadth — the count Share-of-Voice tools optimize — matched the AI’s actual pick in only 1 of 6 categories, and runner-ups often sat on more lists than the picks they lost to. So don’t optimize toward closing a breadth gap; do use it to confirm you exist in the conversation at all.
What should I measure instead?
Attribute-Tied Prose Share: freeze your category’s defining attribute from a neutral “how to choose” guide, then track your share of the independent-editorial sentences that tie that attribute to a brand. It ranked the AI’s pick #1 in 4 of 6 categories in our study — four times Share of Voice’s rate.
If framing lost to frequency, is “narrative ownership” even real?
It’s real but partly misnamed. The signal that predicts is largely the volume of attribute-tied prose, not how cleverly you’re framed — raw count matched the pick in 5 of 6 categories versus the framing-centrality measure’s 3 of 6. Framing (centrality) still decides the low-breadth upsets like Mullvad, where 82% of mentions tie the brand to “privacy.” So: earn volume of the right prose first; framing is the tiebreaker, not the plan.
Can I just buy this on one big placement?
No. Engines read almost entirely different source URLs (about 0.5% overlap) yet converge on the same pick, so the only thing that travels between them is the attribute-language itself, repeated across many independent sources. The move is consistent attribute-tied prose across community, editorial and reviews — not one perfect page.
Does this prove I can move the AI’s pick by earning prose?
No — and we say so plainly. This week tested prediction, not causation. Attribute-tied prose predicts the pick; whether a challenger can manufacture it and shift the recommendation is next week’s experiment.
Sources & method:
- GeoParrot GEO Lab exp7 — 18 brands across 6 buyer-intent categories, re-scored for narrative ownership on the same frozen independent corpus (no engine re-query). Full data and charts in Thursday’s results; pre-registered method, rubric and four predictions in Tuesday’s method post; the opening fact-check in Monday’s post.
- GEO Lab — We Scored the AI Consensus Pick Against Its Runner-Ups (the frozen 18-brand roster, breadth data, and the 1-of-6 Share-of-Voice baseline this verdict rules against).
- GEO Lab — AI Engines Cite Different Sources — But Do They Agree? (the ~0.5% source overlap behind Rule 0: language travels between engines, links don’t).
- GEO Lab — Is Reddit Still a Top AI Citation Source? (the community-citation weighting behind the Perplexity row).
- GEO Lab — Can You Reverse-Engineer the AI Consensus Pick? The Verdict (no single counting signal separates the pick; prose reputation named as the lever we scored this week).
- Context — Best GEO Tools 2026, the Share-of-Voice-centric tooling landscape this week stress-tested.

Leave a Reply