Can You Score “Narrative Ownership” — and Does It Out-Predict Share of Voice? Our Pre-Registered Method

Quick answer: Monday we showed that Share of Voice ranks the AI’s runner-up ahead of the winner — the consensus pick appeared on fewer best-of lists (82.6% vs 93.2%), had a lower rating (4.18★ vs 4.42★) and fewer reviews, yet still won the recommendation. The lever we named was narrative ownership: owning the language of your category’s defining attribute in the independent text engines read. But “own the narrative” is a slogan until you can score it. This is the method, published before the numbers. Three parts, each frozen before we code a single mention: (1) freeze each category’s defining attribute from a neutral “how to choose” source — before we look at which brand owns it, so we can’t retrofit; (2) score attribute-shorthand ownership per brand in the same independent corpus Share of Voice counts — how often the brand is named as the shorthand for that attribute, and what share of the category’s attribute-language it owns; and (3) run it head-to-head against Share of Voice — which ranking’s #1 matches the AI’s actual consensus pick more often. We’ve locked four predictions before scoring a single brand, on the same 18-brand roster where Share of Voice already failed.

Yesterday’s fact-check opened this week’s question: the metric the entire 2026 GEO tooling wave optimizes — Share of Voice — is a floor, not a ranking signal. It told us what doesn’t separate the pick. It also named a candidate for what does: narrative ownership. Today, before we score a single brand, here’s exactly how we’ll test whether that candidate is real and measurable — with the same pre-registration discipline we used for our Reddit causation method and our cross-engine overlap method.

What question are we actually answering?

One testable claim, stated plainly:

Hypothesis (H1): “Narrative ownership” — how much a brand owns the shorthand language of its category’s defining attribute in independent prose — is a measurable signal, and it separates the AI consensus pick from its runner-up more cleanly than any count-based Share-of-Voice signal. Picks own more of their category’s defining-attribute language than the runner-ups they beat, even in the categories where the pick sits on fewer lists — and when narrative ownership and Share of Voice disagree about who’s #1, narrative ownership sides with the brand the engines actually recommend.
Null hypothesis (H0): Narrative ownership is either too subjective to score reliably (two coders can’t agree what “owns the attribute” means), or, once scored, it tracks the pick no better than Share of Voice — the pick doesn’t systematically own more attribute-language than the runner-up, and raw mention counts explain the recommendation just as well. If H0 holds, “own the narrative” is an unfalsifiable slogan and we should say so.

If H1 holds, the GEO scorecard gets a new top line: not “how many lists mention you,” but “do independent writers reach for you as the shorthand for the thing your category is bought for.” If H0 holds, we will have tested our own Monday hypothesis against the outcome it’s supposed to explain — instead of assuming it’s true because it sounds right. The method below is built to tell those two worlds apart, and it’s built to fail loudly if we’re wrong.

How do you turn “narrative ownership” into a number?

This is the whole experiment, so we freeze the definitions before we score anything. The trap with a soft concept like “narrative” is that you can always find a story where the pick looks like the owner. So we split scoring into three locked steps, and the first one exists specifically to stop us from cheating:

Step Definition (frozen before scoring) How we measure it
1 · Freeze the defining attribute The one buying criterion the “best [category]” query turns on — the concept a category is reached for. Chosen from a neutral, independent “how to choose a [category]” framing, never from the pick Before scoring any brand, we log one attribute per category plus the independent source it came from (e.g. VPN → “privacy / no-logs / anonymity”). The full frozen attribute table ships with the results
2 · Attribute-association rate Of the brand’s mentions in independent prose, the share that tie the brand to that attribute in the same clause or sentence — “Mullvad, the no-logs choice,” not a bare list entry Two independent coders classify each brand mention in the corpus as attribute-linked or generic, against a locked rubric; we report the rate and the coders’ disagreement rate
3 · Ownership share Of all the category’s defining-attribute shorthand (“the private one,” “best for privacy,” “built for anonymity”), the share that points to this brand — share of the language that matters, not share of raw mentions We tally every attribute-shorthand phrase in the corpus and attribute it to a brand (or “none”); each brand’s ownership share is its count over the category total

Steps 2 and 3 are deliberately different lenses. Attribute-association rate asks “when this brand comes up, is it framed as the answer to the category’s core question?” — a per-brand centrality. Ownership share asks “who owns the concept overall?” — a category-level, winner-take-share view that is Share of Voice’s structure (a share of a total) applied to the language that carries intent instead of to every raw mention. Reporting both is the point: if a brand wins on share only by being mentioned more, that’s Share of Voice wearing a new hat, and we’ll catch it in step 3 of the test below.

Why freeze the attribute first — and from a neutral source?

Because this is where a “narrative” study goes to die. If you pick the attribute after looking at the pick, every pick owns its narrative by construction — you just describe the pick and call it the category’s essence. So the attribute is frozen before any brand is scored, and it’s derived from independent “how to choose a [category]” guidance — the same kind of neutral editorial that feeds the corpus — not from what the winner happens to be good at. We publish the attribute and its source for all six categories, so a reader can object “that’s the wrong attribute for accounting” and re-run our scoring against their own. The one attribute we’ve already disclosed is VPN → privacy/no-logs, because Monday’s Mullvad case made it public; the other five are frozen and ship Thursday. If the picks only own their attributes once we let ourselves choose the attributes, that’s not narrative ownership — that’s storytelling, and the method is designed to expose it.

What roster and corpus are we scoring — and why reuse them?

We reuse the exact frozen roster and corpus from last week’s experiment (exp6), and the reuse is the honest choice, not a shortcut. The roster is 18 brands across 6 buyer-intent categories — help desk, live chat, marketing automation, survey, accounting and VPN — each already labeled consensus pick, runner-up or challenger by two engines that independently agreed on the #1. Two reasons to hold it fixed:

  • The outcome is already frozen, so we can’t fish. The label — which brand the engines recommend — is the hard, tempting-to-retrofit part, and it was locked last week before we ever thought about narrative. We’re not re-running engines until the recommendations flatter our new metric.
  • It makes the comparison apples-to-apples. We score narrative ownership on the same independent corpus Share of Voice was counted on — the ~10 frozen best-of and comparison lists per category, plus the review-platform prose. The point of the week is to re-read the exact text Share of Voice reduced to a link count and ask what the language says instead. Different corpus, and the two posts wouldn’t line up.

Where the best-of lists are too terse to carry attribute language (a bare “1. Mullvad” with no prose), we supplement with a small, pre-listed set of independent editorial “best [category] for [attribute]” articles, flagged as such in the roster so the two source types can be read separately. The tradeoff, stated up front, is the same one we always state: this is a small, hand-coded sample — treat every result as “the pattern is strong,” not “these percentages to the decimal.”

What are the three parts of the test?

Part What it measures What “narrative ownership is real” looks like
1 · Pick vs runner-up ownership Whether picks own more attribute-language than the runner-ups they beat Picks out-own their runner-up in most categories — including the ones where the pick sits on fewer lists
2 · The disagreement cases Where Share of Voice and narrative ownership rank brands differently, which one sides with the pick In the SoV-failure categories (Mullvad-type), narrative ownership ranks the pick #1 where SoV ranks it last
3 · Head-to-head prediction Which metric’s #1 matches the AI’s actual pick more often across all 6 categories Narrative ownership’s top brand = the pick in more categories than Share of Voice’s top brand does

Part 1 — Pick vs runner-up ownership

For every category we compare the pick’s attribute-association rate and ownership share against its labeled runner-up’s. The clean win for H1 isn’t “picks have high ownership” — strong brands generally do. It’s the gap holding up in the categories where Share of Voice said the pick should lose: the ones where the pick appears on fewer independent lists than the runner-up. If narrative ownership is the real lever, it should be present exactly where the count-based signal is absent. If picks and runner-ups own their attribute-language equally, H1 is already in trouble.

Part 2 — The disagreement cases

The most informative rows are where the two metrics point opposite directions. Mullvad is the archetype: last in the category on Share of Voice (1 of 8 lists, 176 reviews, 3.5★) but — by hypothesis — first on narrative ownership of “privacy.” We catalogue every category where the SoV rank and the narrative-ownership rank disagree by two or more places, and record which metric’s ordering matches the engines’ recommendation. A metric that only agrees with Share of Voice when Share of Voice is already right adds nothing; the value is entirely in the disagreements.

Part 3 — Head-to-head prediction

The decisive test, and the one that can kill H1 outright. On the frozen roster we produce two rankings per category — one by Share of Voice (breadth and mention count, straight from exp6) and one by narrative ownership — and ask a single question: whose #1 is the brand the engines actually recommend? The bar is already public and low: last week Share of Voice’s own breadth leader was the pick in just 1 of 6 categories. Narrative ownership earns its keep only if it clears that bar by a clear margin. And there’s a built-in trap we’re pre-committing to check: if narrative-ownership share only out-predicts because it correlates with raw mention volume, then it’s Share of Voice re-labeled, and we’ll report the attribute-association rate — the volume-independent centrality measure — as the tiebreaker. If rate tracks the pick and raw count doesn’t, the mechanism is framing, not frequency. If neither beats Share of Voice, H1 is dead and Friday’s verdict says so.

What did we lock before collecting a single number?

Pre-registration means the predictions are on the record now, before we score anything, so we can’t quietly move the goalposts on Thursday:

  1. Picks out-own their runner-ups. The consensus pick has higher attribute-shorthand ownership than its labeled runner-up in at least 4 of 6 categories — and the gap survives in the categories where the pick sits on fewer independent lists than the runner-up.
  2. The Mullvad inversion holds. In the categories where Share of Voice and narrative ownership most disagree, narrative ownership ranks the AI’s pick at or near the top while Share of Voice ranks it at the bottom.
  3. Narrative ownership out-predicts Share of Voice. Its top-ranked brand matches the engines’ actual pick in strictly more categories than Share of Voice’s top-ranked brand does (the 1-of-6 bar from exp6).
  4. Framing beats frequency. Attribute-association rate (centrality) predicts the pick better than raw attribute-mention count (volume). If ownership only wins by re-counting mentions, we declare it a re-skinned Share of Voice — not a new signal.

If the data contradicts these, we publish the contradiction. A method that can only confirm itself isn’t a method — it’s a press release.

Where could this break? (The honest limitations)

  • Coding is more subjective than counting. Deciding “is this mention attribute-shorthand or generic?” is a judgment Share of Voice’s list-counting never has to make — and that’s a real cost, not a footnote. We mitigate with a locked rubric, two independent coders, and a published inter-coder disagreement rate; if coders can’t agree, that is the finding (H0’s “too subjective to score”).
  • Retrofitting the attribute is the fatal failure mode. Choose the attribute the pick already owns and you’ve proven nothing. The freeze-first-from-a-neutral-source step is the guardrail, and publishing the attribute table lets readers audit it — but we’re naming the risk openly because it’s the one that would quietly invalidate the whole thing.
  • Correlational, not causal. Even if picks own more narrative, ownership might be a byproduct of genuinely being the best tool, not a lever a challenger can pull. This week only tests whether it predicts; Friday’s verdict takes up whether it’s actionable.
  • Small, hand-run sample. Six categories, 18 brands, coded by hand. Strong patterns, not decimal precision — and the same engine-throttle caveats we’ve flagged before carry over from the labels we reused.
  • Reused roster cuts both ways. Holding the roster fixed makes us comparable to the Share-of-Voice numbers we’re trying to beat, but it also means any quirk in last week’s 6-category draw rides along. We’d want a second, independent category set before calling this a law rather than a strong signal.

FAQ

What is “narrative ownership” in GEO?
It’s how much a brand owns the language of its category’s defining attribute in the independent text AI engines read — how often writers reach for it as the shorthand for the thing the category is bought for (“the private VPN,” “the all-in-one platform”). It’s the qualitative lever we identified when Share of Voice failed to explain the AI’s pick; this week we test whether it can be scored like a metric.

How is narrative ownership different from Share of Voice?
Share of Voice counts mentions — your share of all the times a brand comes up. Narrative ownership counts attribute-language — your share of the specific shorthand that carries the category’s buying intent, and how central that framing is when you’re mentioned. A brand can have low Share of Voice (few lists, like Mullvad’s 1 of 8) but high narrative ownership (it owns “privacy”). Which one predicts the engines’ pick is exactly what this experiment measures.

Isn’t scoring “narrative” hopelessly subjective?
It’s more subjective than counting lists — we don’t pretend otherwise. That’s why the attribute is frozen from a neutral source before any brand is scored, two independent coders apply a locked rubric, and we publish the disagreement rate and the full coded roster. If two careful coders can’t agree what “owns the attribute” means, that failure is a result — it would confirm the null hypothesis that the concept can’t be measured.

Why reuse last week’s brands instead of a fresh study?
Because the outcome — which brand the engines recommend — was frozen last week, before we thought about narrative, so we can’t retrofit it. And scoring on the same independent corpus Share of Voice was counted on is the only way to make the comparison apples-to-apples: we’re re-reading the exact text a SoV tool reduced to a link count and asking what the language says instead.

When do the results come out?
We score the roster midweek and publish the full data and the head-to-head on Thursday, then rule on Friday: does narrative ownership out-predict Share of Voice, or not? The predictions above are locked now so you can hold us to them — including the four ways we’ve committed to declaring ourselves wrong.

Sources

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *