How We’re Testing Whether One GEO Strategy Works Across AI Engines (Our Method)

Quick answer: To test whether one GEO strategy can win across every AI engine, we’re asking 4 engines (ChatGPT Search, Perplexity, Google AI Mode, Gemini) the same 15 buyer-intent questions, logging every cited domain, and measuring the overlap between engines with a Jaccard score. We’ve pre-registered our predictions before collecting any data: if a universal strategy existed, the engines would cite mostly the same sources (high overlap). We predict the opposite — under 15% of sources shared by all four. This post is the full method, published before the results so you can hold us to it.

Yesterday we argued there is no single GEO strategy that wins across every AI engine — backed by a 680-million-citation industry study showing just 11% domain overlap between ChatGPT and Perplexity. Fair criticism: that’s someone else’s data, on B2B SaaS, in their conditions.

So this week at GEO Lab we’re running it ourselves — bigger than our first 3-engine test, and fully in the open. Here’s exactly how, before we touch the data.

What question are we actually answering?

One testable claim, stated plainly:

Hypothesis (H1): AI engines cite mostly different sources for the same question, so optimizing for one engine does not transfer to the others.
Null hypothesis (H0): Engines cite largely the same sources — which would mean a single, universal GEO strategy is possible.

If H0 is true, we should see high overlap: the same domains appearing across ChatGPT, Perplexity, Google AI Mode, and Gemini. If H1 is true, overlap will be low and most cited sources will be unique to one engine. We let the numbers decide.

Which engines and questions are we using?

Four engines — the ones that actually matter for 2026 buyers, including the two we left out of our first test (ChatGPT Search and Gemini):

  • ChatGPT Search (web-search mode on)
  • Perplexity (default web answer)
  • Google AI Mode (post-I/O-2026 default)
  • Gemini

And 15 commercial-intent questions — the kind real buyers ask, where getting cited actually drives revenue. We deliberately avoid navigational or trivia queries. Examples from the set:

  • “What are the best GEO tools for 2026?”
  • “How do I get my brand cited by ChatGPT?”
  • “Perplexity vs ChatGPT for research — which is better?”
  • “Best B2B SaaS analytics platforms”
  • “How to optimize a website for AI search”

How do we collect the data without fooling ourselves?

Nondeterminism and personalization are the two ways a citation experiment quietly lies to you. Our controls:

Risk How we control for it
Personalized results Logged-out / fresh sessions, no memory, neutral account, same region
Run-to-run variance Each question asked 3 times per engine; we record all citations and report variance, not a single lucky pull
Time drift All 4 engines queried within the same 48-hour window
Wording bias Identical question text across every engine — no per-engine rephrasing
Collection error We capture the raw cited URLs the engine shows, then reduce to root domains

Collection is done with a browser agent driving each engine’s real interface — the same approach as our first experiment, where we learned ChatGPT needs its web-search toggle explicitly on and Perplexity throttles after ~5 rapid queries. We pace accordingly.

What exactly do we measure?

Every cited link becomes a root domain (e.g. reddit.com, g2.com). Then we compute:

  • All-engine overlap: % of domains cited by all four engines — the true “universal source” number.
  • Pairwise Jaccard: overlap between each engine pair (e.g. ChatGPT ∩ Perplexity ÷ their union), so we can compare to the industry’s 11% figure.
  • Single-engine share: % of all cited domains that appear on only one engine.
  • Reddit dependence: Reddit’s share of citations per engine (the industry says ~46% Perplexity vs ~0.1% Gemini — does ours match?).
  • Source-type mix: community vs. vendor vs. publisher vs. review-site, per engine.

Our predictions — locked in before the data

Pre-registration is what separates an experiment from a story told after the fact. So here’s what we expect, on the record, before collecting a single citation:

  1. Under 15% of domains will be cited by all four engines.
  2. Over 70% of cited domains will be unique to a single engine.
  3. Reddit will be heavy on Perplexity and near-zero on Google AI Mode / Gemini.
  4. ChatGPT will cite the fewest distinct domains (most conservative).

If we’re wrong, we’ll say so. That’s the point of writing the predictions down where you can see them.

Next at GEO Lab: we collect the citations over the next two days and publish the raw results later this week — every number against these four predictions. Then on Friday we deliver the verdict: is “no universal GEO strategy” confirmed by our own data, or do we have to walk it back? Start with the claim →

FAQ

Why pre-register predictions?
Because it’s easy to look at results and craft a tidy explanation afterward. Writing the predictions first makes the test falsifiable — if the data contradicts us, you’ll know we didn’t move the goalposts.

Why only 15 questions and 4 engines?
Small enough to run rigorously and transparently by hand, large enough to show a pattern. We’d rather publish 15 questions we actually controlled than 500 we can’t vouch for. A focused test that matches a 680M-citation study is a real signal.

How do you handle AI answers changing every time?
We ask each question three times per engine and report the spread, so a single odd response can’t drive the conclusion.

Can I replicate this myself?
Yes — that’s the goal. The questions, engines, and metrics are all listed above. Run it on your own niche and you’ll likely see the same per-engine divergence.

Sources

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *