Quick answer — Verdict: the consensus pick is reverse-engineerable, but the lever isn’t any signal you can count. This week we pre-registered four predictions about which reputation signal separates an AI “consensus” pick from an equally-good runner-up, then scored 18 brands across 6 buyer-intent categories. Three of the four failed. Best-of breadth doesn’t separate the pick — runner-ups sat on as many or more independent lists in 5 of 6 categories (picks averaged 82.6% of a category’s lists, runner-ups 93.2%). Star rating is a floor everyone clears (all three groups in a 4.2–4.5★ band; the picks average the lowest, 4.18★). And Mullvad — the AI’s #1 VPN on just 1 of 8 lists with 176 reviews — breaks the “breadth floor” outright. But the pick is not a coin flip: prominence (top-3 placement) trended in the right order, and every pick was the one brand independent writing treats as the category’s obvious answer. So the honest verdict on Monday’s claim is conditional: reputation is reverse-engineerable, but the lever is category-defining prose reputation, not the breadth, rating, or review volume the 2026 playbooks tell you to grind. Below: the scored predictions, a platform-by-platform checklist, and next week’s question.
This closes week five of the GEO Lab. On Monday we fact-checked the claim that a brand becomes the AI consensus pick when independent third parties converge on it — and you can’t crown yourself. Tuesday we pre-registered a method: score three candidate signals — best-of breadth, review rating/volume, and challenger break-in — against a frozen roster of picks vs runners-up, with four predictions locked before we measured a single brand. Thursday we published the raw result: no single countable signal separates the pick. Today we turn it into a verdict on the week’s claim and a checklist you can run Monday.
So what’s the verdict on this week’s claim?
The hypothesis we locked Tuesday was strong and specific: the consensus pick is separated from an equally-capable runner-up by third-party reputation breadth more than by any signal a brand controls. The data refutes it. Breadth didn’t separate the pick — it ran slightly higher for runner-ups. There was no breadth floor. The one countable signal that even trended right, top-3 prominence, showed an 8-point pick-vs-runner gap that overlaps far too much to be “the separator.” Here’s how the four pre-registered predictions scored:
| Pre-registered prediction | Verdict | What the data said |
|---|---|---|
| 1 · Breadth is the dominant separator (picks on ≥2× the lists; a floor runners-up don’t cross) | ❌ Refuted | Picks averaged 82.6% of lists vs runner-ups’ 93.2%; runner ≥ pick in 5 of 6 categories; no floor (Mullvad on 1 of 8) |
| 2 · Rating is a floor; volume & recency are the real lever | ⚠️ Partly | Rating-as-floor held (all in 4.2–4.5★, pick lowest at 4.18★). But volume didn’t separate (median and mean disagree) and recency was login-walled — n/a |
| 3 · Self-published volume ≈ 0 as a lever | ✅ Confirmed | Consistent with exp5: the #1 brand’s own site was the cited source in just 4.2% of answers — self-publishing doesn’t move the pick |
| 4 · The moat is breadth-gated — no low-breadth pick | ❌ Refuted | Mullvad is a clean counterexample: the AI’s #1 VPN on 1 of 8 lists, 176 reviews, 3.5★ |
Verdict on H1: refuted as stated — but H0 doesn’t win either. Breadth is not the dominant lever. But the null hypothesis — “picks and runners-up are identical on every signal, so the pick is a coin flip you can’t target” — is also wrong. The pick is predictable; it’s just predictable from something the rubric couldn’t count. Final verdict: conditional yes. You can reverse-engineer the consensus pick, but only if you stop measuring placements and start measuring narrative.
If breadth doesn’t decide it, what does?
Every countable signal we tested is a floor, not a separator. Breadth, rating, and review volume are table stakes: a brand that appears on almost no lists, sits below ~4.2★, or has a thin review profile rarely enters the consensus set at all. But once a handful of strong brands have all cleared those floors — which is exactly the pick-vs-runner situation — the countable signals stop discriminating. Our runner-ups already matched or beat the picks on breadth, rating, and often volume. Adding a ninth listicle or a thousand more 4.4★ reviews doesn’t flip the engine to naming you first, because your rivals are already there.
What separates the pick is whether the brand is the category’s defining narrative — the answer independent prose reaches for by default. Mullvad is the proof. It is both engines’ identical #1 VPN, yet it sits on 1 of 8 best-of lists, carries 176 Trustpilot reviews at 3.5★, and has no meaningful G2 or Capterra presence — while runner-up ExpressVPN (7 of 8 lists, ~28k reviews) and challenger Surfshark (8 of 8 lists, ~31k reviews) dominate every countable signal. If you reverse-engineered the pick from breadth and reviews, you would never land on Mullvad. The engines do, because Mullvad owns the category’s story — no-logs, RAM-only, anonymous accounts, “the privacy purist’s VPN.” That prose was written across independent forums and reviews long before anyone counted a list, and the model absorbed it. This is the same mechanism we found in our earlier citation work: prose reputation lives in community writing and independent media, and the domains engines reach for are the ones that carry that writing — not your own site.
Does that mean you can’t reverse-engineer the pick?
No — and this is the useful half of the verdict. “Reverse-engineerable” doesn’t require a metric you can dashboard. It requires knowing what to earn. The failure of breadth/rating/volume is good news for a challenger: it means you can’t be locked out simply because an incumbent has 30,000 more reviews than you. Mullvad wasn’t. What you have to earn instead is ownership of a sharp, defensible slice of the category’s narrative — a claim so specific that independent writers use your brand as the shorthand for it. That’s slower and less checklist-able than racking up placements, which is precisely why it’s a moat and not a growth hack.
What should you actually do? (the platform-by-platform checklist)
Since the separator is prose reputation, the work is earning it on the surfaces AI engines demonstrably read — not on your own site. Run this in order; the top three are the earned layer, the last is the tuning layer.
- Independent review & best-of media (the table-stakes floor). Get on the ~10 independent lists your category’s engines actually cite, and aim for top-3 prominence, not just presence — prominence was the only countable signal that trended with the pick. Then stop. A ninth list won’t separate you; treat this as a threshold to clear, not a number to maximize.
- Community & forums — Reddit, niche communities, Stack-style Q&A (where the narrative is written). This is where category-defining prose is actually authored. Your goal isn’t a link — it’s to become the answer people type unprompted when someone asks “best [category] for [specific need].” Own one specific use-case (“the privacy one,” “the one for solo founders,” “the cheapest that still does X”) so the shorthand attaches to you. Do this the way we’ve documented — earned mentions, never self-promotion, which fails.
- Your own category-defining content (the anchor, not the win). You can’t self-rank into the pick — Monday’s fact-check and prediction 3 both confirm it. But you can publish the definitive, first-data resource on your narrow slice so independent writers cite you as the source of the story. The move is authoring the category’s reference material, not writing a “Top 10” that ranks yourself #1.
- Per-engine tuning (the 30% on top — do it last). Only after the earned layer is in place. Our own data shows engines agree far less than they overlap on citations, so don’t assume the pick transfers. Sanity-check your target query on each engine (ChatGPT with web search on, Perplexity, Google AI Mode, Gemini) and use the per-engine checklist for surface-specific tweaks. Win the shared, reputation-driven layer first; treat per-engine work as the finish, not the foundation.
What are the honest limits of this verdict?
This is a separation test on a fixed outcome, not an intervention — we can show a signal tracks (or fails to track) the pick, but we can’t prove any signal causes it, and we won’t. The roster is 18 brands across 6 categories; the patterns are strong but not decimal-precise, and one category (VPN) drives the most dramatic counterexample. Recency was genuinely unmeasurable — every review platform hides the date distribution behind a login wall — so “review freshness” remains an open sub-question, not a closed one. And “prose reputation” is, by nature, harder to quantify than a placement count; naming it as the separator is a hypothesis the countable signals failed to displace, not a variable we measured directly. Turning it into something measurable is exactly next week’s job.
What’s the one-line rule to take away?
Clear the countable floors — lists, rating, reviews — then stop grinding them and go earn the one sentence the category tells about you. The AI consensus pick is the brand that owns a story, not the brand with the biggest numbers.
What are we testing next week?
If “prose reputation” is the real lever, the obvious challenge is: can we measure it? Next week’s question — can a brand’s category-narrative ownership be quantified from independent text, and does that score predict the consensus pick better than breadth did? We’ll try to build a rough “narrative-ownership” signal from the language independent sources use about a brand (how often it’s named as the shorthand for a specific attribute) and score it against the same frozen roster. If a text-based narrative score separates picks from runners-up where breadth failed, we’ll have turned this week’s qualitative verdict into something you can actually target. That’s the seed for Monday.
Frequently asked questions
What’s the final verdict — can you reverse-engineer the AI consensus pick?
Conditionally, yes. Our locked hypothesis that best-of breadth is the dominant separator was refuted (3 of 4 predictions failed), but the pick isn’t random either. It’s predictable from category-defining prose reputation — being the brand independent writing treats as the obvious answer — which no countable signal (breadth, rating, review volume) captures. You can target it; you just can’t dashboard it with placement counts.
Should I stop trying to get on best-of lists and collect reviews?
No — clear those floors, because brands below them rarely enter the consensus set. Just don’t expect them to separate you once you’re in the running: our runner-ups already matched or beat the picks on breadth, rating, and often volume. Beyond the floor, the return on a ninth listicle or a thousand more reviews is near zero. Spend that effort earning narrative ownership instead.
How can Mullvad be the #1 VPN with only 176 reviews and 1 of 8 lists?
Because the consensus pick is anchored in prose reputation, not counts. Mullvad is the category’s reference point for no-logs, privacy-purist VPNs, and AI engines absorbed that narrative from independent writing long before anyone tallied a list. Its runner-up and challenger dominate every countable signal and still don’t get the pick.
Does this transfer across AI engines?
Don’t assume so. Our earlier experiments show engines overlap far less than people expect, so win the shared, reputation-driven layer first and treat per-engine tuning as the last 30%. See our per-engine checklist for surface-specific tweaks.
Isn’t “prose reputation” too vague to act on?
It’s less checklist-able than a placement count — that’s the point, and why it’s a moat. The concrete action is to own a sharp, specific slice of the category’s story so independent writers use your brand as the shorthand for it. Making that ownership measurable from text is exactly what we test next week.

Leave a Reply