We Scored the AI Consensus Pick Against Its Runner-Ups — No Single Reputation Signal Separates Them

Quick answer: We scored 18 brands across 6 buyer-intent categories — each labelled a consensus pick, runner-up, or challenger — against the three reputation signals we pre-registered on Tuesday, and no single signal cleanly separates the pick from its runner-up. Best-of breadth doesn’t do it: runner-ups appeared on as many or more independent lists than the AI’s #1 in 5 of 6 categories (picks averaged 82.6% of a category’s lists, runner-ups 93.2%). Star rating doesn’t do it: picks, runners-up and challengers all cluster in a 4.2–4.5★ band, and the picks actually average the lowest (4.18★ vs 4.42★ vs 4.45★). Review volume is muddy — picks lead on the median, trail on the mean. And the cleanest counterexample kills the “breadth floor” idea outright: Mullvad, the AI’s #1 VPN pick, sits on just 1 of 8 independent lists with 176 reviews at 3.5★, while its runner-up and challenger blanket 7 and 8 of those lists. Three of our four locked predictions failed. The thing that separates a consensus pick isn’t a signal you can count — it’s category-specific prose reputation, which none of breadth, rating, or review volume captures.

This is Thursday’s result in this week’s GEO Lab arc. Monday we established the AI consensus pick is real and you can’t self-manufacture it; Tuesday we published the method and four predictions before measuring a single brand. Today we run the numbers against those predictions. As promised, we’re reporting what the data said — including the three predictions it contradicted.

Does best-of breadth separate the pick from the runner-up? (No.)

Prediction 1 was the strong one: consensus picks would appear on at least 2× the independent best-of lists their runners-up do, and every pick would clear a “breadth floor” no runner-up crosses. It failed on both counts. Averaged across categories, picks landed on 82.6% of a category’s independent lists — below the runner-ups’ 93.2%. In 5 of 6 categories the runner-up was on as many or more lists than the AI’s #1; only live chat had the pick strictly ahead.

Grouped bar chart of best-of breadth by category. Consensus pick vs runner-up as share of that category's independent best-of lists. help desk: pick 83, runner 100. live chat: pick 100, runner 71. marketing automation: pick 100, runner 100. survey: pick 100, runner 100. accounting: pick 100, runner 100. VPN: pick 12 (Mullvad, 1 of 8 lists), runner 88. In 5 of 6 categories the runner-up matches or beats the pick on breadth.

There is no floor. If a “majority of lists” were the price of entry, Mullvad — on 1 of 8 VPN lists, named in the top three of exactly zero — could not be the pick. It is. Breadth clearly matters as table stakes (challengers that appear on few lists rarely make the set), but among strong brands it does not pick the winner. One honest nuance: top-3 prominence did trend in the right order — picks were named prominently on 69% of lists, runners-up 61%, challengers 35% — the only signal that even monotonically tracks the label. But the pick-vs-runner gap (8 points, heavily overlapping) is far too small to call it “the separator.”

Is it the star rating, then? (No — it’s a floor everyone clears.)

Prediction 2 had two halves. The first — rating alone doesn’t separate; it’s a floor — held cleanly. Picks, runners-up and challengers all sit inside a tight 4.2–4.5★ band, and the picks average the lowest of the three (4.18★). A brand needs a credible rating to be in the conversation, but once it clears ~4.2★, a higher number does not earn it the #1 spot.

Bar chart of mean review rating by label. Consensus pick 4.18 stars, runner-up 4.42 stars, challenger 4.45 stars, all inside a shaded 4.2 to 4.5 star band. The AI's pick averages the lowest rating of the three groups.

The second half — that review volume and recency would show a wider gap and act as the real lever — did not hold. Volume was muddy: picks led on the median review count (5,328 vs runners’ 3,815) but trailed on the mean (8,088 vs 8,861), because a couple of enormous runner-up/challenger profiles (ExpressVPN ~28k, Surfshark ~31k Trustpilot reviews) skew the average. No clean pick-vs-runner separation either way. And recency — the freshness sub-signal — we simply could not measure: every major review platform hides the review-date distribution behind a login wall. We flagged that risk in the method; it came true, and we’re not going to fake a number to fill the gap.

Can a brand be the pick without the breadth? (Yes — meet Mullvad.)

Prediction 4 said the moat is breadth-gated: no brand would sit in the consensus set without first clearing the breadth floor — no shortcut, no counterexample. The data produced the counterexample on the first VPN query. Plot every brand by its best-of breadth and its review count and the picks (navy) scatter across the entire field — there is no region of the chart that is “where picks live.”

Scatter plot of best-of breadth (share of category's independent lists, x-axis) versus review count (log scale, y-axis) for 18 brands, colored by label. Consensus picks in navy are scattered from the bottom-left to top-right of the chart, showing no single region isolates the pick. Mullvad, the AI's number 1 VPN pick, sits alone in the bottom-left corner at 1 of 8 lists and 176 reviews at 3.5 stars, while its runner-up and challenger sit at high breadth with tens of thousands of reviews.

Mullvad is the clean break. It is the AI’s identical #1 VPN pick across both engines, yet it has no meaningful G2 or Capterra presence, 176 Trustpilot reviews, and a 3.5★ average — while runner-up ExpressVPN (7 of 8 lists, ~28k reviews) and challenger Surfshark (8 of 8 lists, ~31k reviews at 4.3★) dominate every countable signal. If you were reverse-engineering the pick from breadth and reviews, you would never land on Mullvad. The engines do, because Mullvad is the category’s privacy-purist reference point — the story reviewers and forums tell about it (no-logs, RAM-only, anonymous accounts) is the consensus, and the model absorbed that prose long before it counted any list. That’s a reputation you can’t reduce to a placement count.

How did the four pre-registered predictions score?

We wrote these down before measuring anything. Here’s the honest scorecard.

Pre-registered prediction Verdict What the data said
1. Breadth is the dominant separator (picks on ≥2× the lists; a floor runners-up don’t cross) ❌ Refuted Picks averaged 82.6% of lists vs runners-up’s 93.2%; runner ≥ pick in 5 of 6 categories; no floor (Mullvad 1 of 8)
2. Rating is a floor; volume & recency are the lever ⚠️ Partly Rating-as-floor confirmed (all in 4.2–4.5★, pick lowest). But volume didn’t separate (median vs mean disagree) and recency was login-walled — n/a
3. Self-controlled signals ~zero association with being the pick ✅ Held Consistent with exp5’s finding that just 4.2% of #1 picks traced to the brand’s own site (cross-referenced, not freshly measured this week)
4. The moat is breadth-gated, no low-breadth pick ❌ Refuted Mullvad is a clean low-breadth pick: AI’s #1 VPN on 1 of 8 lists, 176 reviews

Three of four wrong. Our null hypothesis — that no countable third-party signal cleanly separates picks from runners-up — is largely what survived. But not because the pick is a coin flip, as the null assumed. The separator exists; it just isn’t a number you can rack up.

What are the honest limits here?

This is 6 categories and 18 brands — directional, not a census. Specific caveats we won’t paper over:

  • Lists were 6–8 per category, not the 10 we planned. TechRadar, PCMag, Forbes and other classic best-of domains hard-block bots (403), so we froze the independent lists we could actually read. Fewer lists widens the error bars on every breadth number.
  • Recency is untested. The freshness sub-signal from Prediction 2 is genuinely unmeasured, not measured-and-null — login walls, full stop.
  • Some review counts are snippet-derived. G2/Trustpilot often 403 the direct page, so a few counts come from search snippets; QuickBooks’ US (1.1★) vs UK (3.9★) Trustpilot split is unstable enough that we leaned on G2 for it.
  • Labels are frozen from exp5 (Gemini + Perplexity, two engines) — we did not re-query the engines this week, by design, so the picks are as stable as that 2-engine consensus.

None of these rescue the three failed predictions — if anything, thinner list coverage should have helped a breadth-floor show up, and it still didn’t.

What does this mean if you’re doing GEO?

The popular 2026 advice — “get on more best-of lists, rack up reviews” — is directionally fine as table stakes and useless as a differentiator. Our runner-ups already match or beat the picks on both. If you’re a strong-but-not-picked brand, adding a 9th listicle or a thousand more 4.4★ reviews is not what flips the AI to naming you first. What separates the pick is whether you are the category’s defining narrative — the brand the independent prose treats as the obvious answer, the way Mullvad owns “private VPN.” That’s slower and less checklist-able than a placement count, which is exactly why it’s a moat. Friday we turn this into the verdict on Monday’s claim and a concrete reputation-signal playbook: given that the separator is prose reputation, what should a challenger actually spend the next quarter earning?

Frequently asked questions

Did any single reputation signal separate consensus picks from runners-up?
No. Across 6 buyer-intent categories, best-of breadth ran slightly higher for runners-up (93.2% vs 82.6% of lists), star rating was a 4.2–4.5★ floor all three groups cleared (picks lowest at 4.18★), and review volume disagreed with itself (picks led the median, trailed the mean). Three of our four pre-registered predictions failed.

How can Mullvad be the AI’s #1 VPN with only 176 reviews?
Because the consensus pick is anchored in category-defining prose reputation, not review counts. Mullvad is the reference point for private, no-logs VPNs, and AI engines absorbed that narrative from independent writing — it appears on just 1 of 8 best-of lists and carries a 3.5★, 176-review Trustpilot profile, yet still wins the pick over ExpressVPN (~28k reviews) and Surfshark (~31k reviews).

Does this mean best-of lists and reviews don’t matter for GEO?
They matter as table stakes — challengers absent from the lists rarely enter the consensus set at all. They just don’t separate the eventual pick from an equally-listed, equally-rated runner-up. Treat breadth and rating as thresholds to clear, not levers to out-pull a rival on.

Why couldn’t you measure review recency?
Every major review platform (G2, Capterra, Trustpilot) hides the distribution of review dates behind a login wall, so the “share of reviews in the last ~6 months” sub-signal we pre-registered was unmeasurable. We flagged that risk in the method; rather than fabricate a figure, we report it as n/a.


Sources & method:

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *