The Verdict: Our Geometry Rule Failed Its Own Blind Test at 4/6 — So We Rewrite It, Not Defend It. Monday’s Bigger Claim (GEO Findings Explain, They Don’t Predict) Just Got Proven on Us

Quick answer: Our geometry rule failed its own blind test — 4 of 6 — and by the bands we locked on Tuesday (5/6 promotes, 4/6 inconclusive, ≤3/6 retracts), 4/6 is a fail, not a pass. So we neither celebrate nor recant — we rewrite. The verdict has three parts. The binary appearance version of the rule is ❌ refuted: a neutral “best [category] in 2026?” returns a segmented listicle (“best for automation,” “best free,” “best for enterprise…”), so almost every credible tool appears, and “did the challenger show up?” is too easy a bar — both our same-axis HELD calls (ActiveCampaign, Pipedrive) missed on it. What survives is the ✅ distinct-axis direction: every challenger we said would cross on an axis the incumbent structurally can’t claim — Brave, Plausible, Obsidian, Bitwarden — appeared in all three engines (4/4 categories, 12/12 engine-observations), and a post-hoc re-score at “best-overall crown” resolution sorts all six categories cleanly (same-axis incumbents Mailchimp and Salesforce lost the crown 0/3, both to distinct-axis outsiders). That is a 🟡 conditional result — promising, but post-hoc, so it is a hypothesis for next week’s pre-registration, not a confirmed rule, and we label it that way. On Monday’s bigger claim — that most 2026 GEO findings explain the past instead of predicting the future — the verdict is a clean ✅ confirmed: our own rule fit last week’s data 6-for-6 in hindsight and still couldn’t clear the out-of-sample bar. Retrodiction is not prediction, and we just proved it on ourselves.

This is Friday’s verdict, closing this week’s GEO Lab arc. Monday we argued that the month’s impressive GEO studies — 680M citations analyzed, a 46x cross-engine brand-citation gap — almost all explain data already collected rather than predict an unseen case, and that our own geometry-gate rule had exactly that weakness. Tuesday we pre-registered six blind predictions across six fresh categories and fixed the pass/fail line in public. Wednesday we queried the engines blind. Thursday we scored: 4 of 6, with the distinct-axis arm perfect and both same-axis arms broken. Today we rule — and turn the wreckage into something you can use.

What exactly are we ruling on?

Two claims of different sizes were on the line this week, and the honest thing is to grade them separately:

  • The specific rule. “A challenger crosses the AI-search incumbency moat when it owns a distinct attribute axis the incumbent structurally can’t claim; it stays held when it’s a better/cheaper/simpler version of the incumbent’s own axis.” Last week it fit 6 of 6 categories — but only in hindsight.
  • The meta-claim. “Most 2026 GEO findings explain the past instead of predicting the future — including ours — so any rule you’d bet a budget on has to survive an out-of-sample, pre-registered, blind test.” That was Monday’s real thesis.

The elegant, uncomfortable part is that testing the first is the cleanest possible test of the second. If our own carefully built rule couldn’t predict blind, that is the strongest evidence Monday’s meta-claim is right — even about us.

Did the geometry rule pass its own pre-registered test? ❌ No — 4/6 is a fail by our own bands

We committed the bands before collection: 5 or 6 correct promotes the rule to a working (still provisional) predictor; 4 is inconclusive — reported as not passing; 3 or fewer forces a retraction. The blind test landed on 4 of 6. That is the middle band. We do not get to say “the rule works,” because it didn’t clear the line we drew — and we do not have to say “the rule is overfit garbage,” because it stayed well above the coin-flip floor. The intellectually honest word is fail, in the specific sense that it failed to earn the promotion it was tested for.

And the failure was not random. All four distinct-axis predictions hit; both same-axis predictions missed. A coin flip does not sort its errors perfectly by the exact variable you coded on — so the miss is telling us where the rule is wrong, not just that it’s unproven.

So is Monday’s bigger claim confirmed? ✅ Yes — we proved it on ourselves

This is the cleanest verdict of the week. Monday’s thesis was that a rule fit to data you already have will look great on that data and tell you little about the next case — and that the GEO field routinely sells the first as if it were the second. Our geometry rule was the perfect test subject: it fit 6 of 6 categories last week, a flawless retrodiction. If retrodiction reliably became prediction, a 6-for-6 hindsight fit should have cruised past a 5-of-6 blind bar. It didn’t — it dropped to 4/6 the moment we removed our foreknowledge of the answers.

That drop is the finding. A perfect fit on data you already had is the single least trustworthy result in empirical work, and we just watched our own perfect fit lose a point under the only test that counts. So the next time you read a GEO study that “found” the trait winning brands share — including anything we publish — ask the one question that separates a lever from a story: was it ever stated in advance and checked against a case it hadn’t seen? If not, treat it as an explanation. We built the whole llms.txt myth-check on the same reflex, and this week we aimed it at ourselves.

Does any of the rule survive? ✅ The direction does — 4/4 — and a crown re-score is promising but unproven

Yes, and it’s worth being precise about which part. The directional claim — own an axis the leader structurally can’t claim and the AI will surface you — went a clean 4 for 4 on brand-new categories: Brave (privacy), Plausible (cookieless), Obsidian (local-first), and Bitwarden (open-source) each appeared in the recommendation set of all three engines. That is blind, out-of-sample, and reproduced on categories the rule had never touched. It is not the retrodiction we flagged last week.

What broke was the binary event definition. We had defined “held” as “the challenger does not appear,” and in a 2026 AI answer that is almost impossible — the answer is a segmented listicle where nearly every credible tool gets a “best for ___” slot. So both same-axis challengers “crossed” on appearance and both HELD calls missed. But when you re-score the same frozen data at a sharper, more decision-relevant event — who did the engine crown “best overall”? — the geometry signal comes back, and it sorts all six:

Two-panel verdict scorecard. Left panel, the pre-registered appearance test: 6 categories, the four distinct-axis CROSS calls (Brave, Plausible, Obsidian, Bitwarden) hit in coral, the two same-axis HELD calls (ActiveCampaign, Pipedrive) missed in grey, final tally 4 of 6 which fails the 5-of-6 promotion bar. Right panel, the post-hoc best-overall crown re-score: in all 6 categories the brand holding or taking the crown owns a distinct axis (coral), and neither same-axis challenger ever takes the crown while both same-axis incumbents Mailchimp and Salesforce lose it 0 of 3 to distinct-axis outsiders HubSpot and Klaviyo — 6 of 6 directionally, but marked post-hoc and unproven.
  • Distinct-axis challengers took or shared the top: Obsidian was crowned best overall in 2 of 3 note-taking answers; Brave was the practical/top pick in some browser answers; where the incumbent kept the crown (Google Analytics 3/3, 1Password 2/3), the distinct-axis challenger still owned its named slot without needing to dethrone anyone.
  • Same-axis challengers never took the crown: ActiveCampaign got “best for automation,” Pipedrive got “best for sales pipelines” — a slot, never the crown.
  • Same-axis incumbents lost the crown outright: Mailchimp kept “best overall” in 0 of 3 (demoted to “beginners / simple newsletters” while HubSpot and Klaviyo took the top); Salesforce kept it in 0 of 3 (all three crowned HubSpot). And the brand that dethroned them was itself a distinct-axis outsider — HubSpot’s free all-in-one ecosystem, Klaviyo’s ecommerce depth — the same reversal we’ve watched all thread, where a brand wins on what it owns rather than raw volume.

The honest caveat, stated loudly: this crown re-score is post-hoc. We changed the yardstick after seeing the results. That is exactly the move we spent Monday warning against, so it earns zero credit as confirmation. It is a hypothesis — a sharper, more testable one — and the only way it earns the word “rule” is to be pre-registered and run blind, which is next week’s job.

What do we now believe — stated as the rewritten rule?

Compressing the week into what we’d actually stake a claim on today:

  • ✅ Confirmed (blind, 4/4): A challenger that owns a distinct axis the incumbent structurally can’t claim gets surfaced by AI search — it will appear in the recommendation set on a generic query.
  • ❌ Refuted: “Appearance” is not “winning.” In a segmented listicle, appearing is nearly automatic for any known tool, so the old binary CROSS-vs-HELD test measures almost nothing.
  • 🟡 Conditional / next to test: The decision-relevant event is which slot you own and whether you hold “best overall.” Geometry appears to predict that too (6/6 post-hoc), but that must be earned blind before we believe it.

Net verdict on the specific rule: directionally real, mechanically rewritten, still on probation. Net verdict on the meta-claim: confirmed.

What does this mean platform by platform? (the part you can act on today)

The generic query behaves differently on each engine, so the checklist differs. Across all three, one thing holds: appearing is table stakes; owning a named slot is what a distinct axis buys you — and there is no single universal move.

ChatGPT (web search on) — win the lead sentence

ChatGPT tends to open with a single lead pick and a one-line reason (“it would be Brave … blocks ads and trackers by default”), and it is the stingiest engine on brand citations (~0.59% in the 34,234-response benchmark). So being the named distinct pick matters more here than anywhere.

  • Own one attribute the category leader structurally can’t claim, and state it as your first-sentence identity — not a feature list.
  • Get third-party prose (Reddit threads, review sites) repeating that exact attribute phrase; ChatGPT rewards a crisp, corroborated one-liner.
  • Do not position as “a better [incumbent]” — same-axis framing gets you a slot at best and nothing at worst.

Perplexity — own the “best for ___” callout, not the crown

Perplexity cites brands roughly 13% of the time (a 46x gap over ChatGPT) and routinely names a best-overall plus a practical/segmented pick with the axis attached (“Brave, with stronger default privacy”). You don’t need the crown here — you need the axis callout.

  • Ensure one clean, citable source (docs page or comparison) that states your distinct axis in plain language.
  • Frame content around the buyer segment your axis serves (“privacy-first,” “open-source,” “local-first”) so it maps to a named slot.
  • See the engine-specific mechanics in how to rank in Perplexity.

Google AI Mode — inclusion is automatic, so target the row label

Google AI Mode returns the most explicitly segmented listicle — “best for ecommerce / best free / best for enterprise.” Almost every credible tool appears, so mere inclusion is worthless as a goal.

  • Target owning the label of a row (“best free,” “best privacy-first”), not appearing somewhere in the table.
  • Structured comparison content and schema help AI Mode file you under the right named slot — see optimizing for Google AI Mode.
  • Same-axis “better version” content dissolves into the crowd here faster than on any other engine.

The 30-second challenger self-test

Before you spend a dollar on GEO content, answer one question — the same one that sorted this week’s data: does your product own an axis the category leader structurally can’t claim, or are you pitching a better/cheaper version of the leader’s own axis? Distinct axis → invest in owning the named slot for it. Same axis → expect to appear but not to win, and reconsider the positioning before the content. This mirrors our per-engine checklist and the cross-engine reality that engines cite different sources yet converge on a pick.

What are the limits of this verdict?

  • Small and binary. n=6, one call each, one logged-out run per engine. A structured 4/6 with a perfect arm split is suggestive, not decisive.
  • The crown re-score is post-hoc. Worth repeating: it earns zero confirmation credit until pre-registered. We report it as a lead, not a result.
  • Selection judgment. We chose the categories and coded the axes. The public codes and the two required same-axis cases are guardrails, not a full defense.
  • Snapshot and personalization. Engines change weekly and personalize; a pass would describe this snapshot, not a law.
  • Prediction is not mechanism. Even at the crown resolution, geometry would forecast the pick — not prove that owning a distinct axis causes the AI to choose it. The causal claim stays out of scope.

What’s next?

The crown hypothesis is too good to leave post-hoc. Next week we pre-register it properly: fresh categories again, but this time the prediction is who holds “best overall” — distinct-axis brand takes or keeps the crown, same-axis challenger never takes it — with the pass/fail line fixed before we query. If geometry calls the crown blind, the rule finally graduates. If it doesn’t, the appearance-level failure was the whole story and we say so. Either way you’ll read it here first — the same deal we made when we put our own finding on the line.

Frequently asked questions

Did the geometry rule pass or fail?

It failed the specific test we set. We pre-registered a 5-of-6 promotion bar; it scored 4/6, which by our own committed bands is inconclusive and reported as not passing. It did not hit the ≤3/6 retraction line, so we’re not calling last week’s classification overfit garbage — but 4/6 does not earn the word “rule,” and we won’t pretend it does.

If it failed, why aren’t you retracting it entirely?

Because the failure is structured, not random: all four distinct-axis calls hit and both same-axis calls missed, and the two misses share one mechanism (listicle segmentation makes “appears” trivially easy). That tells us the event definition was wrong, not necessarily the direction. Retracting the direction would ignore a 4/4 blind result; defending the binary rule would ignore a real miss. Rewriting it is the honest middle.

Isn’t the “best-overall crown” re-score just moving the goalposts?

It would be, if we claimed it as a win — so we don’t. Changing the yardstick after seeing results is exactly the retrodiction we warned about Monday, so the crown re-score gets zero confirmation credit and is labeled a post-hoc hypothesis. It only becomes evidence when it’s pre-registered and run blind, which is next week.

What’s the single takeaway for a brand doing GEO right now?

Own an axis the category leader structurally can’t claim — privacy, open-source, local-first, cookieless — rather than pitching a “better version” of the leader’s strength. Appearing in the AI’s slate is nearly automatic for any known tool; owning a named slot (or the crown) is what a distinct axis buys you. If your positioning is “cheaper/better [incumbent],” fix the positioning before you write a word of GEO content.

Does this contradict last week’s incumbency-moat verdict?

It sharpens it. Last week’s claim — the moat is a geometry gate, not a size threshold — was retrodictive and we said so. This week’s blind test confirms the direction (distinct axis gets you surfaced) while showing the original binary event was too coarse. The size-doesn’t-decide-it finding stands; the “crossed vs held” definition gets upgraded to “owns a slot vs owns the crown.”

Why publish a failed test at all?

Because a rule put at risk and reported honestly — pass or fail — is worth more than a folder of after-the-fact explanations that were never falsifiable. That’s the entire reason this Lab exists: to test what AI actually recommends, and to be the first to say when our own finding doesn’t survive.

Sources

Related guide: A practical checklist for getting surfaced by Gemini and Google’s AI answers — read how to rank in Google Gemini.

Related guide: See exactly which visitors arrive from ChatGPT, Perplexity, and Google AI Mode — read how to track AI search traffic in GA4.

Related: This experiment is part of our ongoing GEO research. See every headline finding in one place: the 2026 GEO Benchmark.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *