Can Our Geometry Rule Call an AI’s Pick Before We Look? We Pre-Register 6 Blind Predictions — Locked Here, Before a Single Engine Is Queried

Quick answer: Today is a pre-registration, not a result. Last week our geometry-gate rule fit six software categories 6-for-6 — but only in hindsight, because we coded the categories after we already knew which brand the AI picked. So this week we hold the rule to the bar we set on Monday: out-of-sample, blind, pre-registered, falsifiable. Here we lock six predictions across six fresh categories — none of them in last week’s frozen roster — coded before we query a single engine. Four predictions say the challenger crosses the moat (Brave, Plausible, Obsidian, Bitwarden — each owns a distinct axis the incumbent structurally can’t claim); two say the incumbent is held (ActiveCampaign and Pipedrive are “better versions” of the incumbent’s own axis). Wednesday we query the engines blind and score on Thursday. The rule graduates to a working predictor only if it calls at least 5 of 6 correctly; at 3 of 6 or worse — chance for a coin-flip prediction — we retract the geometry gate to a hindsight-fit story, in public, in the same place we made the claim.

This is Tuesday in this week’s GEO Lab arc, and the whole point is discomfort. On Monday we argued that most 2026 GEO “findings” explain the past instead of predicting the future — and that our own geometry finding had exactly that weakness. Listing a limitation is cheap. The only honest follow-through is to make the rule predict something it has never seen, with the answer written down first. That’s what a pre-registration is: a public, timestamped commitment that removes our freedom to reinterpret the result after it lands.

What is the geometry rule, stated as a mechanical prediction?

Last week’s finding, compressed: whether an AI’s challenger pick crosses the incumbency moat depends on the geometry of its story, not the incumbent’s size. To make it testable, we restate it as a rule that produces one binary call per category, with no wiggle room:

  • Distinct-axis → CROSS. If the leading challenger’s primary narrative axis is an attribute the incumbent structurally cannot claim without abandoning its core identity or business model, the rule predicts the AI will name the challenger among its recommended options.
  • Same-axis → HELD. If the challenger’s axis is a “better / cheaper / simpler version” of the incumbent’s own core attribute, the rule predicts the incumbent (or a same-axis giant) leads and the challenger is not in the recommended set.

The size of the incumbent is explicitly not an input. That’s the part last week’s data supported hardest — Mullvad owns “privacy” as the #1 VPN pick despite roughly 160x fewer reviews than the runner-up — and it’s the part most likely to embarrass us if it was a fluke. A blind test is the only way to find out.

How did we pick the categories — and what makes this “blind”?

We fixed a selection rule before choosing anything, so we couldn’t cherry-pick easy wins:

  • Fresh. Six categories, none of which appeared in last week’s frozen 18-brand roster (help desk, accounting, live chat, marketing automation, survey, VPN). These are cases the rule has never touched.
  • Clear structure. Each category has a recognizable market incumbent (the category’s default answer) and one prominent challenger that GEO practitioners would name.
  • Balanced by design, not by convenience. We required at least two same-axis cases, so the test can fail in both directions — a distinct-axis challenger that fails to appear is a miss, and a same-axis challenger that unexpectedly tops the incumbent is also a miss.

“Blind” means the geometry coding below was written from general category knowledge — what each brand is known for — not from the AI’s picks. We have not queried ChatGPT, Perplexity, or Google AI Mode for “best [category]” on any of these six. We obviously know these brands’ reputations; what we do not know is which ones each engine will actually put in its recommended set this week. That gap between “reputation we know” and “AI output we haven’t seen” is exactly what the test measures.

The six pre-registered predictions (locked)

This table is the frozen artifact. Each row was coded before collection; the one-line rationale is logged now so it can’t be rewritten later.

Category Incumbent Challenger Challenger’s primary axis Geometry code Prediction
Web browser Chrome Brave Privacy — blocks ads & trackers by default Distinct (Chrome is ad-funded; can’t make ad-blocking its identity) CROSS
Web analytics Google Analytics Plausible Privacy-first, cookieless, GDPR-friendly, lightweight Distinct (GA’s model is data collection for Google’s ads) CROSS
Note-taking app Notion Obsidian Local-first, plain-text Markdown, you own your files Distinct (Notion is cloud-native by design) CROSS
Password manager 1Password Bitwarden Open-source and free Distinct (1Password is closed-source, subscription) CROSS
Email marketing Mailchimp ActiveCampaign More powerful automation — same job, more power Same-axis (a “better Mailchimp,” not a new axis) HELD
CRM Salesforce Pipedrive Simpler, visual sales pipeline for small teams Same-axis (easier version of the CRM job) HELD

Four CROSS, two HELD. If the geometry rule is real, the four distinct-axis challengers should earn a place in the engines’ recommended set even on a generic query, and the two same-axis challengers should not displace the incumbent. Where the coding was genuinely arguable — pCloud-style cases where a challenger could be read on a second, more distinct axis — we deliberately did not include them, to keep every code a clean call. The judgment we did make is now public and unchangeable.

How will Wednesday’s blind collection work?

To keep the outcome uncontaminated, the protocol is fixed here, in advance:

  • One neutral query per category. “What is the best [category] in 2026?” — asked verbatim, logged in full. The neutral question is deliberately the hard bar: a distinct-axis challenger has to earn the recommended set without us framing the question around its axis.
  • Three engines. ChatGPT (web search on), Perplexity, and Google AI Mode — the same three we’ve used since our cross-engine consensus work, logged out, single run.
  • “Recommended set” defined now. A brand counts as recommended if the engine names it as a top option — a “best overall,” a shortlist, or a ranked top-three. A passing mention (“others include…”) does not count.
  • Scoring rule. CROSS observed = the challenger is in the recommended set in at least 2 of 3 engines. HELD observed = the challenger is not in the recommended set in at least 2 of 3 engines, and the incumbent (or a same-axis giant) is named as a top pick. A prediction is correct when observed matches predicted.
  • No touching the codes. Wednesday’s collector records raw engine output only; it does not see or edit this table. A secondary axis-framed query (“best privacy browser,” etc.) is collected for color, but the pre-registered test is the neutral query alone.

What result would prove the rule wrong?

The pass/fail line is committed now, so Thursday’s write-up can’t move it:

  • 5 or 6 of 6 correct → the geometry gate graduates from a hindsight story to a working (still provisional) out-of-sample predictor.
  • 4 of 6 → inconclusive. Suggestive, not confirmed — reported as not passing.
  • 3 of 6 or worse → we retract. Three correct is exactly what a coin flip yields on six binary calls, so at that level we say plainly that last week’s 6-for-6 fit was overfit, and we downgrade the geometry gate to “a hindsight-fit story” in public.

Concretely, the rule loses if a distinct-axis challenger fails to appear — Bitwarden absent from “best password manager,” or Plausible absent from “best web analytics” — or if a same-axis challenger unexpectedly tops the incumbent, such as Pipedrive being named the best CRM over Salesforce. We are naming those failure modes now, before we can see which way they break.

What are the limits of even a passing result?

A clean pass would be real evidence, but we won’t oversell it:

  • Small and binary. n=6, one call each. A 6-for-6 out-of-sample result is far stronger than a 6-for-6 retrodiction, but it is still six data points — encouraging, not decisive.
  • Selection judgment. We chose the categories and challengers. A skeptic can argue we picked distinct-axis cases that were easy to call; the two required same-axis cases and the public codes are our guardrails, not a full defense.
  • Coding judgment. Distinct vs same-axis is a human call. Locking each code and its rationale publicly, before results, is the only real check — which is why we did it here rather than in a private spreadsheet.
  • Snapshot and personalization. Engines change weekly and personalize; this is a single logged-out run. A pass says the rule predicted this snapshot, not that it holds forever.
  • Prediction is not mechanism. Even 6-for-6, this shows geometry forecasts the pick — not that owning a distinct axis causes the AI to choose it. The causal claim stays out of scope, as it did in our earnability verdict.

That’s the honest frame for the week. We took the same standard we applied to every other GEO finding and pointed it at our own. If geometry calls these six blind, it earns the word “rule.” If it doesn’t, you’ll read that here on Friday — the same place we make every other claim.

Frequently asked questions

Why pre-register instead of just running the test quietly?
Because a private test lets you rationalize afterward — drop an awkward category, redefine “recommended,” reinterpret a near-miss. Publishing the predictions and the pass/fail line first removes those degrees of freedom. If we fudge on Thursday, the receipt is right here.

Isn’t picking the categories yourselves a way to rig the result?
It’s the biggest risk, which is why we fixed a selection rule (fresh categories, clear incumbent-plus-challenger) and required at least two same-axis cases that must go the incumbent’s way. It’s not a perfect defense — an independent category list would be stronger — but the codes are locked and the failure modes are named.

What’s the difference between this and last week’s 6-for-6?
Last week we coded categories after seeing the AI’s picks — the labels were public and the fit was guaranteed to look good. This week the predictions are written before any engine is queried. Same rule, opposite epistemic footing: retrodiction then, prediction now.

Why the neutral “best [category]” query instead of an axis-framed one?
Because the axis-framed version (“best privacy browser”) stacks the deck for a distinct-axis challenger. The neutral question is the harder, fairer test: the challenger has to reach the recommended set on the generic query, which is what a real buyer would ask.

How can I run a version of this on my own category?
Before you query anything, write down one prediction: does your challenger own an axis the incumbent structurally can’t claim (predict it gets named) or a “better version” of the incumbent’s axis (predict it doesn’t)? Then query the engine and check. One honest out-of-sample call beats a folder full of after-the-fact explanations — the whole argument of Monday’s piece.

Sources

Related guide: A practical checklist for getting surfaced by Gemini and Google’s AI answers — read how to rank in Google Gemini.

Related guide: See exactly which visitors arrive from ChatGPT, Perplexity, and Google AI Mode — read how to track AI search traffic in GA4.

Related: This experiment is part of our ongoing GEO research. See every headline finding in one place: the 2026 GEO Benchmark.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *