Most 2026 GEO “Findings” Explain the Past — They Don’t Predict. This Week We Test Whether Our Own Geometry Rule Can Call an AI’s Pick Before We Look

Quick answer: Mostly no — and that’s the honest read of a very good month for GEO data. The August 2026 studies are real: one analyzed roughly 680 million AI citations and found only about 11% of domains are cited by both ChatGPT and Perplexity; another measured 34,234 AI responses and found a 46x gap in how often each engine cites brands (ChatGPT ~0.59%, Perplexity ~13.05%). But nearly all of it is retrodictive — it explains patterns in data already collected, correlationally, after the fact. Explaining what already happened is not the same as predicting what hasn’t, and the GEO field routinely presents the first as if it were the second. Our own geometry-gate rule from last week has exactly this weakness, and we flagged it: the category labels were public when we coded them, the fit was 6-for-6 after the fact, n=6. So this week we hold ourselves to the higher bar. We pre-register the prediction, then query brand-new categories blind — coding each category’s narrative geometry before we see the AI’s pick. If geometry calls the picks out-of-sample, it graduates from a hindsight story to a working rule. If it doesn’t, we say plainly that the earlier fit was overfit.

This is Monday’s fact-check, and it opens this week’s GEO Lab question: can any GEO rule actually predict a brand’s AI pick — before you look — or does it only explain picks you’ve already seen? It’s a meta-question about the whole discipline, and it’s uncomfortable because the honest answer implicates our own past work as much as anyone else’s. That’s the point. Last week we argued the AI-search incumbency moat is a geometry gate, not a size threshold — and it fit our data perfectly. Perfect fit on data you already have is the single least trustworthy result in empirical work. So before we let ourselves believe it, we test it the only way that counts.

What did this month’s big GEO studies actually find?

Give the optimists their due: the data got a lot better in 2026. The headline studies this month are genuinely large and genuinely useful:

  • Engines cite almost different webs. An analysis of roughly 680 million citations reports that only about 11% of domains are cited by both ChatGPT and Perplexity — each engine draws from a largely separate source pool.
  • Brand-citation rates differ by an order of magnitude. A B2B benchmark of 34,234 AI responses found ChatGPT cited brands about 0.59% of the time versus Perplexity’s 13.05% — roughly a 46x spread across engines for the same questions.
  • The audience is now enormous. ChatGPT reportedly grew from ~400M to ~900M weekly active users between early 2025 and early 2026, and Perplexity processes on the order of 100M queries a month — so whatever the citation logic is, it’s deciding a lot of buying attention.

None of this is wrong, and we’ve reported adjacent findings ourselves — that engines cite almost entirely different source URLs and still converge on a pick, and that there is no single universal GEO strategy across engines. The data is real. The question is what kind of claim it licenses.

Why is “explains” not the same as “predicts”?

Every study above is a measurement of data that already exists. It tells you what the citations were. That’s valuable, but it’s retrodiction — fitting an account to observations already in hand. The failure mode is famous in every empirical field: a pattern you find after looking at the data will fit that data almost by construction, because you had the answers in front of you while forming the rule. The test of a rule is whether it holds on cases it has never seen.

Two specific traps show up constantly in GEO content:

  • Correlation dressed as a lever. “Pages with schema get cited more” or “brands with more Reddit mentions rank higher” are correlations in a snapshot. They may reflect a cause, a common third factor, or reverse causation. A correlation predicts nothing about your specific page until it’s been tested as an intervention.
  • Hindsight fit (HARKing). You observe the winners, then name the trait they share, then present that trait as the rule that “explains” the win. With enough traits to choose from, something always separates the winners in the sample. That’s not a discovery; it’s curve-fitting with words. We wrote a whole piece on how a similar retrofitting made llms.txt look more effective than the data supports.

The tell is simple: a retrodictive claim can always show you why the past makes sense; it can rarely tell you what the next case will be. If a GEO “rule” has never been stated in advance and then checked against an unseen result, treat it as an explanation, not a prediction.

Doesn’t our own geometry rule have exactly this problem?

Yes — and we’re not going to pretend otherwise. Last week we scored six software categories and found the AI’s challenger picks split cleanly on narrative geometry: a challenger crosses the incumbency moat when it owns a distinct attribute axis (Mullvad owning “privacy” as the #1 VPN despite ~160x fewer reviews than the runner-up), and stays stuck when its story is just a “better version” of the incumbent’s own axis (Freshdesk vs Zendesk, Xero vs QuickBooks). The rule fit 6 of 6 categories, and the size story drew no line at all.

But we coded those categories after we already knew which brands the AI picked. The labels were public. The sample was six. A 6-for-6 fit under those conditions is exactly the retrodiction we just warned about — a very good hindsight story, not a demonstrated predictor. We said so in the verdict itself, and listing that limitation isn’t enough. The only way to know whether geometry is a rule or a rationalization is to make it predict something it hasn’t seen.

What would a real predictive test of a GEO rule look like?

Four properties turn an explanation into a test. This is the bar we’re holding ourselves to this week:

  • Out-of-sample. Brand-new categories that were not in last week’s frozen roster — cases the rule has never touched.
  • Blind. Code each category’s geometry (is the leading challenger on a distinct axis, or a better-version of the incumbent’s axis?) before querying any engine — so the coding can’t be nudged by knowing the answer.
  • Pre-registered. Write the specific prediction for each category down first, publicly, with the pass/fail condition fixed in advance. No moving the goalposts after results land.
  • Falsifiable. State what result would prove the rule wrong. If the geometry call misses on the unseen categories, the rule fails — and we report that as the finding.

Almost none of the month’s GEO studies clear this bar, and that’s not a knock on them — measurement and prediction are different jobs. But if you’re going to act on a GEO rule — bet a content budget on it — you want a rule that has predicted, not just explained.

What are we testing this week?

The rest of the week is one experiment, run in the open:

  • Tuesday — pre-register. Pick a set of fresh categories, code each one’s narrative geometry blind, and write down the predicted AI pick and the pass/fail line before any engine is queried.
  • Wednesday — collect blind. Query the engines on the unseen categories and record the actual picks, without touching the pre-registered codes.
  • Thursday — results. Score prediction against outcome, with charts.
  • Friday — verdict. If geometry called the picks out-of-sample, it earns the word “rule.” If it missed, we retract it to “a hindsight-fit story” — publicly, in the same place we made the claim.

Either way you get a result that most GEO content never gives you: a rule that was put at risk before the data came in. That’s the whole reason this Lab exists — to test what AI actually recommends, and to be the first to say when our own finding doesn’t survive.

Frequently asked questions

Are the August 2026 GEO studies unreliable?
No — they’re useful measurements. The 680M-citation overlap and the 46x cross-engine gap are real and worth knowing. The caution is narrow: a measurement of past citations describes what happened; it doesn’t by itself predict your next result. Use them to understand the landscape, not as forecasts for a specific page.

What’s the difference between retrodiction and prediction?
Retrodiction fits an explanation to data you already have; prediction states an outcome for a case you haven’t seen yet and is then checked against it. A rule can be great at the first and useless at the second — which is why out-of-sample testing is the only real test.

Why are you testing your own finding instead of defending it?
Because a perfect fit on data we already had is the least trustworthy result in empirical work. If the geometry gate is real, a blind test will confirm it. If it isn’t, we’d rather find out ourselves than mislead readers who might act on it.

What is “narrative geometry” in one line?
Whether a challenger owns a distinct attribute axis the incumbent can’t claim (crosses the moat) or is pitching a better version of the incumbent’s own axis (stays stuck). Last week it fit 6 of 6 categories — in hindsight. This week we see if it predicts.

How can I apply this to my own GEO decisions now?
Before you act on any GEO “rule,” ask one question: has it ever been stated in advance and then checked against a case it hadn’t seen? If not, treat it as an explanation and test it small on your own category before betting a budget on it.

Sources

  • SEO Sherpa — AI Search Statistics 2026 (680M-citation cross-engine overlap; engine usage scale).
  • Averi.ai — ChatGPT vs Perplexity vs Google AI Mode: B2B SaaS Citation Benchmarks (2026) (34,234 responses; 0.59% vs 13.05% brand-citation rates).
  • Omnibound — AI Search Statistics (2025–2026): 55+ Data Points on GEO, Citation Rates.
  • GeoParrot GEO Lab first-party data — incumbency-moat verdict, cross-engine consensus results, no universal GEO strategy.

Free tool: Generate a valid, structured llms.txt for your site in seconds with our free llms.txt Generator — or browse all of GeoParrot’s free GEO & utility tools.

Related guide: A practical checklist for getting surfaced by Gemini and Google’s AI answers — read how to rank in Google Gemini.

Related: This experiment is part of our ongoing GEO research. See every headline finding in one place: the 2026 GEO Benchmark.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *