Quick answer: The 2026 GEO Benchmark is GeoParrot’s running dataset of pre-registered, first-party experiments on how AI search engines choose what to cite. The headline findings: across four engines and fifteen buyer-intent questions, only 0.5% of cited domains overlapped across all four; Reddit appears in 67% of Google AI answers but is the linked source only 4% of the time; and in the largest llms.txt study (137,210 domains) AI search bots accounted for just 1.1% of fetches. The through-line: citation is decided per-question, not by domain size.
Last updated: August 2026. This is a living benchmark — we add each new GEO Lab experiment as it publishes. Cite any figure with a link to this page.
Why this benchmark exists
Most GEO advice is asserted, not tested. We run a public GEO Lab: we pre-register a hypothesis, query real engines, freeze the data, and publish the result — hits and misses alike. This page collects the headline number from each experiment in one place so you can cite it, sanity-check vendor claims against first-party data, and prioritize the tactics that actually move citations.
The benchmark at a glance
| Question we tested | Headline finding | Full experiment |
|---|---|---|
| Do AI engines cite the same sources? | No. 4 engines × 15 questions, 386 domains — only 0.5% cited by all four | Cross-engine citations |
| Do 15 domains control AI citations? | Claimed 68%; our buyer-intent data says top 15 = 23%, 77% cited once | Citation concentration |
| Does a Reddit mention cause a citation? | Contributing signal, not a lever — Reddit in 67% of Google AI answers, linked 4% | Reddit causation |
| Is llms.txt worth it? | 97% of files get zero requests; AI bots = 1.1% of fetches (137,210 domains) | llms.txt data |
| Can you self-rank into a recommendation? | No — a big vendor’s own listicle was cited in only 8.3% of 24 answers | Self-ranking trap |
| Can you reverse-engineer the consensus pick? | Yes, but 3 of 4 predictions failed — best-of breadth doesn’t separate it | Consensus signals |
| Is Share of Voice the right metric? | Wrong as a ranking signal — optimize attribute-tied prose instead | Share of Voice |
| When does a challenger beat an incumbent? | The moat is a geometry gate (distinct axis), not a size threshold | Incumbency moat |
| Can our own rule predict picks blind? | Partly — it scored 4/6 on out-of-sample blind predictions | Blind test |
| Does AI search use backlinks? | Not like PageRank — links shape citations indirectly via authority | Backlinks |
1. There is no single “AI” to optimize for
We asked four AI engines the same fifteen buyer-intent questions. They surfaced 386 cited domains — and just 0.5% were cited by all four. The overwhelming majority were unique to one engine. Practical takeaway: GEO is per-engine work, so use the per-engine checklist rather than a one-size approach. Full method and numbers in the cross-engine results and why there’s no universal GEO strategy.
2. Citation concentration is real but overstated
A widely-shared index of 680M citations claimed 15 domains capture 68% of AI citation share. On our own first-party, buyer-intent data the top 15 domains accounted for 23%, across 220 distinct domains, with 77% cited only once. Concentration exists, but the long tail is very much alive — a focused site can still get cited. See the concentration re-run.
3. Reddit is a contributing signal, not a lever
Reddit shows up in 67% of Google AI answers, but it’s the actual linked source only 4% of the time. So a Reddit mention correlates with — and probably feeds — citation, but posting to Reddit is not a direct citation lever. Treat community presence as trust-building, not a shortcut. Details in the Reddit causation verdict and is Reddit still AI’s #1 source?
4. llms.txt is mostly ignored today
In the largest study yet (137,210 domains), 97% of llms.txt files received zero requests, and AI search bots accounted for just 1.1% of fetches — while Google has said it doesn’t use the file. It doesn’t hurt, but it isn’t the shortcut it’s sold as. See the llms.txt data and the complete llms.txt guide.
5. Self-ranking listicles don’t buy citations
The popular advice — publish your own “best-of” list, rank yourself #1 — underperforms. A large vendor’s own listicle was cited in only 8.3% of 24 buyer-intent answers. Engines discount transparently self-serving sources. The full anti-self-ranking playbook is in the self-ranking verdict.
6. Share of Voice is the wrong ranking metric
Much of the 2026 GEO tooling wave optimizes “share of voice.” Our data says that’s the wrong target as a ranking signal — what separates a cited pick is attribute-tied prose (owning a specific, named strength), not breadth of mentions. See the Share of Voice verdict.
7. Consensus picks aren’t separated by “best-of breadth”
We scored 18 brands and pre-registered four predictions about what separates an AI “consensus” pick from an equally-good runner-up. Three of four failed: breadth of best-of appearances didn’t distinguish the winner. What did correlate was owning a distinct attribute. Read the consensus verdict.
8. The incumbency moat is a geometry gate, not a size threshold
Across a frozen six-category roster, whether a challenger could beat an incumbent’s default didn’t depend on the incumbent’s size. It depended on whether the challenger owned a distinct attribute axis — a geometry gate. Being “a better version of the same thing” didn’t cross the moat. See the incumbency-moat verdict.
9. We hold our own rule to the same standard
To avoid hindsight bias, we pre-registered six blind, out-of-sample predictions using our geometry rule and scored them before looking. It hit 4 of 6 — good, not gospel — so we rewrote the rule rather than defending it. Falsifiable, dogfooded GEO is the whole point. See the blind-test verdict.
How we run these experiments
- Pre-registration: we write the hypothesis and predictions before querying any engine.
- Real queries: we ask live engines (ChatGPT, Perplexity, Google AI Mode, and others) real buyer-intent questions.
- Frozen data: we capture the answer prose and cited sources, then score without re-querying.
- Publish hits and misses: failed predictions are reported as prominently as successes.
What to do with these findings
- Optimize per engine — there’s no universal citation list.
- Own a distinct attribute; don’t just be a “better version” of the incumbent.
- Invest in original data and clear expertise over volume or self-ranking tricks.
- Treat llms.txt and share-of-voice dashboards as nice-to-haves, not priorities.
- Measure citations and AI referrals — start with the best GEO tools and GA4 tracking.
Frequently asked questions
Can I cite these numbers?
Yes. Please link to this benchmark page or the specific experiment. Every figure comes from a pre-registered, first-party GeoParrot experiment with its full method published.
How often is the benchmark updated?
It is a living document. We add each new GEO Lab experiment as it publishes and revise the summary table with the latest headline figures.
Why do your numbers differ from other GEO studies?
Two reasons: we test buyer-intent questions specifically, and we use first-party data with a frozen, pre-registered method. Many cited industry figures mix query types or measure appearances rather than links, which inflates them.
What is the single most important takeaway?
Citation is decided per question, not by domain size. Own a distinct, well-evidenced angle and optimize for each engine separately.
The bottom line
The 2026 GEO Benchmark exists to replace assertion with evidence. The consistent signal across ten experiments: AI citation rewards distinct, credible, per-question expertise — not size, self-ranking, or file-based shortcuts. Start from the GEO playbook, then use these numbers to prioritize.

Leave a Reply