The Blind Test Scored 4 of 6: Our Geometry Rule Called Every Distinct-Axis Pick Correctly — and Missed Both “Same-Axis Holds” Because the Challenger Crossed Anyway

Quick answer: We scored the 6 blind predictions we locked Tuesday against fresh queries to ChatGPT, Perplexity, and Google AI Mode — and hit 4 of 6. The distinct-axis arm was flawless: every challenger we said would cross on an axis the incumbent structurally can’t claim — Brave (privacy), Plausible (cookieless), Obsidian (local-first), Bitwarden (open-source) — appeared in the recommendation set of all three engines. That’s 4/4 categories, 12/12 engine-observations. But the same-axis arm missed both: we predicted Mailchimp and Salesforce would hold against a “better version” challenger, and instead ActiveCampaign and Pipedrive crossed in all three engines too. By our pre-registered bands — ≥5 promotes the rule, 4 is inconclusive, ≤3 forces a public retraction — 4/6 lands squarely in the inconclusive zone. The honest read: this isn’t random 4/6. The CROSS arm is bulletproof and the HELD arm collapsed as a unit — because a neutral “best [X] in 2026?” answer is a segmented listicle where nearly every credible tool earns a “best for ___” slot, so mere appearance is too easy a bar. What Friday’s verdict has to settle: is the rule salvageable at a stricter resolution, or is the HELD half dead?

This is Thursday’s result in this week’s GEO Lab arc. Monday we argued that most 2026 GEO “findings” explain the past but don’t predict, and put our own geometry rule on the line; Tuesday we pre-registered 6 predictions and locked them — 4 distinct-axis CROSS, 2 same-axis HELD — before a single engine was queried, precisely to escape the hindsight trap that shadowed last week’s 6-for-6 retrodiction. Today we ran the neutral query once per category across three engines and scored blind against the frozen table. Nothing in the prediction table was edited after collection — that’s the whole point of pre-registration.

How did the blind predictions score?

4 of 6. Four hits, two misses — and the split is not where a coin-flip would put it. Every one of the four distinct-axis calls landed, and both same-axis calls missed. Here’s the frozen scorecard, scored against the recommendation set each engine returned.

Category Challenger (axis) Predicted Observed (of 3 engines) Verdict
Web browser Brave (privacy — distinct) CROSS CROSS — 3/3 ✅ hit
Web analytics Plausible (cookieless — distinct) CROSS CROSS — 3/3 ✅ hit
Note-taking Obsidian (local-first — distinct) CROSS CROSS — 3/3 ✅ hit
Password manager Bitwarden (open-source — distinct) CROSS CROSS — 3/3 ✅ hit
Email marketing ActiveCampaign (automation — same-axis) HELD CROSS — 3/3 ❌ miss
CRM Pipedrive (simpler — same-axis) HELD CROSS — 3/3 ❌ miss
Horizontal bar scorecard of the 6 blind predictions. All six challengers appeared in all 3 engines (bars reach 3 of 3). The four distinct-axis rows (Brave, Plausible, Obsidian, Bitwarden), coloured coral, were predicted CROSS and hit — 4 of 4. The two same-axis rows (ActiveCampaign, Pipedrive), coloured navy, were predicted HELD but observed CROSS — 0 of 2 hit. Final tally 4 of 6.

Read against our locked bands, 4/6 is the inconclusive result — short of the 5/6 that would have promoted the geometry rule to something we trust predictively, but well clear of the ≤3/6 coin-flip line that would have forced us to retract last week’s classification as overfit. We don’t get to celebrate and we don’t have to recant. We do have to explain the shape of the miss, because it’s the opposite of random.

Why did the distinct-axis arm go a perfect 4-for-4?

Because a distinct axis gives the engine a reason to name a second brand that the incumbent structurally can’t absorb. Every one of the four played out the same way in the prose: the incumbent kept a slot, and the challenger got handed the slot the incumbent can’t fill. ChatGPT’s browser answer opened with “it would be Brave … blocks ads and trackers by default”; Perplexity’s own practical pick was Brave “with stronger default privacy” even while it named Chrome best-overall. In analytics, all three engines filed Plausible under an explicit “privacy-first / cookieless” row. In notes, Obsidian owned the “local-first / knowledge-management” slot in every engine — two of them ranked it top overall. In password managers, Bitwarden was the “best free / open-source” pick across the board. Twelve engine-observations, twelve appearances. The geometry claim — own an axis the leader can’t stand on and the AI will surface you — reproduced cleanly on categories that were never in our roster before Tuesday. That’s the genuinely encouraging half of the result, and it’s blind and out-of-sample, so it isn’t the retrodiction we flagged last week.

Why did both “same-axis holds” miss?

Because the bar we set — “challenger appears in the recommendation set” — turns out to be trivially easy to clear in a segmented listicle. A neutral “best email marketing platform in 2026?” doesn’t return a single winner; it returns a table of “best for ecommerce / best for automation / best for creators / best value.” ActiveCampaign was handed the “advanced automation” slot in all three engines. Pipedrive got “best for sales pipelines / sales-focused SMBs” in all three. Neither is a “better Mailchimp” or a “better Salesforce” that the leader crushes; each is a named use-case pick, so each cleared the CROSS bar even though we’d coded it same-axis and predicted a hold. In other words, the miss isn’t that our axis geometry was wrong — it’s that presence in a listicle is not the same event as winning, and our HELD prediction implicitly assumed the incumbent would keep the challenger out. In a table with seven slots, almost nobody gets kept out.

Did the incumbents at least keep the crown?

This is where the same-axis story gets worse for the naive rule and more interesting for the real one. If “appears in the set” is too generous, the sharper event is “best overall” — the crown the engine hands to one brand. On that measure, the categories split hard.

Grouped horizontal bar chart. For every category, the challenger appeared in all 3 engines (coral bars all reach 3). But the incumbent's retention of the 'best overall' crown (navy bars) varies widely: Web analytics GA4 kept it 3 of 3, Password manager 1Password 2 of 3, Web browser Chrome 1 of 3, Note-taking Notion 0 of 3, Email Mailchimp 0 of 3, CRM Salesforce 0 of 3. The two same-axis incumbents lost the crown outright.

Google Analytics held “best overall / the default” in 3 of 3 engines even though Plausible crossed — the distinct-axis challenger took the privacy slot without dislodging the leader. But the two same-axis incumbents lost the crown outright: Mailchimp kept it in 0 of 3 (demoted to the “beginner / simple newsletters” tier while HubSpot and Klaviyo took best-overall), and Salesforce kept it in 0 of 3 (all three engines crowned HubSpot best-overall and left Salesforce with “best for enterprise”). The brand that displaced them wasn’t even our coded challenger — it was a third party owning a distinct axis (HubSpot’s free all-in-one ecosystem, Klaviyo’s ecommerce depth). So the same-axis categories didn’t just fail to “hold” the challenger out; the incumbent got knocked off the top by someone else entirely. That reversal — leader displaced by a distinct-axis outsider — is the same pattern we’ve watched all thread, where a brand can be first on what it owns rather than on raw volume.

So is the geometry rule confirmed, refuted, or something in between?

In between — and specifically located. The directional claim survives every test we threw at it: distinct-axis challengers cross (4/4 here, and the two same-axis incumbents were themselves toppled by distinct-axis outsiders). What failed is the binary CROSS-vs-HELD event definition at the resolution we chose. “Appears in the recommendation set” is the wrong yardstick, because 2026 AI “best-of” answers are structurally pluralistic — they name a slate, not a winner. The rule’s real currency isn’t presence; it’s which slot you own and whether you hold “best overall.” That’s a testable reframing, and it’s exactly what Friday’s verdict has to rule on: do we re-score the same frozen data at the “best-overall” resolution and see whether geometry predicts the crown — or do we accept that the appearance-level HELD prediction is dead and rewrite the rule? Either way, we report it here first, before deciding. That’s the deal we made when we put our own finding on the line.

Frequently asked questions

What exactly was pre-registered, and can you prove it wasn’t edited after collection?

The full prediction table — 6 categories, each coded distinct-axis or same-axis, each mapped to CROSS or HELD, plus the pass/inconclusive/retract bands — was published Tuesday in the method post, before any engine was queried. The scoring protocol was locked in the same post: one neutral query “best [category] in 2026?” per category, three engines, challenger in the recommendation set of ≥2 of 3 engines counts as an observed CROSS. Today’s post scores against that frozen table without changing a cell.

Why does “appears in the recommendation set” undercount as a test?

Because a modern AI answer to “best [category]” is a segmented listicle: it returns “best for ecommerce,” “best for automation,” “best free,” and so on. In a table with six or seven use-case slots, almost every credible tool appears somewhere, so “did the challenger appear?” is a low bar. The sharper, more decision-relevant event is “who did the engine crown best overall?” — and that’s where the same-axis incumbents actually lost.

Does 4/6 mean the rule is wrong?

No — it means it’s unproven at this resolution. Our pre-registered bands call 4/6 inconclusive on purpose: it’s below the 5/6 promotion line and above the ≤3/6 retraction line. The directional claim (distinct axis → cross) went 4/4 and is the part that held; the binary hold-prediction is the part that broke. Friday’s verdict decides what to do about it rather than papering over it.

Could the misses just be noise from one query per engine?

Unlikely to be pure noise, because the failure is structured, not scattered: all four distinct-axis calls hit and both same-axis calls missed, and the two misses share a mechanism (listicle segmentation). Random error wouldn’t line up perfectly with the axis coding. Single-query sampling is a real limitation we’ll name in the verdict, but it doesn’t explain a clean arm-vs-arm split.

What does this mean if I’m a challenger brand doing GEO?

The actionable half is confirmed: to get surfaced by AI search, own an axis the incumbent structurally can’t claim — privacy, open-source, local-first, cookieless — rather than pitching a “better version” of the leader’s own strength. Appearing in the slate is nearly automatic for any known tool; owning a named slot (or the crown) is what a distinct axis buys you. Our per-engine checklist and the no-universal-strategy finding both point the same way.

GeoParrot is a GEO Lab: we pre-register predictions, query the engines, and report what actually happened — including when our own rule only half-lands. Friday we deliver the verdict on whether the geometry rule survives at the “best-overall” resolution or gets rewritten. Methods and raw per-engine captures available on request.

Related guide: A practical checklist for getting surfaced by Gemini and Google’s AI answers — read how to rank in Google Gemini.

Related guide: See exactly which visitors arrive from ChatGPT, Perplexity, and Google AI Mode — read how to track AI search traffic in GA4.

Related: This experiment is part of our ongoing GEO research. See every headline finding in one place: the 2026 GEO Benchmark.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *