Do a Handful of Domains Control AI Citations? The Verdict — Plus a Per-Format GEO Playbook

Quick answer — Verdict: the concentration story is false where GEO competes, and the real common thread is a page format, not a domain. “A handful of domains control everything” holds only for general-knowledge answers. For the buyer-intent queries your business actually fights over, the top 15 domains capture just 23% of citations, spread across 220 distinct domains, with 77% cited only once. What the domains AI engines choose have in common is not a type or an authority tier — it’s a shape: 66% of every citation is a ranked “best-of” listicle or comparison page (92% of the top 25 URLs). Even vendors, who take the largest share, earn 62% of their citations from their own best-of content, not product pages. The lever is clear: get into the best-of lists that already rank, or publish one worth citing. Below: the scored predictions and a per-format playbook.

This closes week three of the GEO Lab. On Monday we fact-checked the “15 domains = 68%” claim against our own buyer-intent data and got 23%. Tuesday we pre-registered a method and locked four predictions. Thursday we published the raw dissection — and hit a reversal: engines choose a page format, not a domain type. Today we turn all of it into a verdict and a checklist.

So do a handful of domains control AI citations?

Verdict: ❌ False for buyer-intent queries — conditionally true only for general knowledge. The 5WPR index that reported 68% concentration across 680M citations isn’t wrong; it’s measuring a different question. Its corpus is dominated by general-knowledge answers where Reddit and Wikipedia genuinely soak up citation share. But when we reran the exact same concentration math on our own 306 citations across 219 domains for commercial “best tool for X” queries, the world inverted:

  • Top 15 domains = 23.1% of citations, not 68%.
  • 220 distinct domains cited across just 12 questions.
  • 170 domains (77.6%) cited exactly once.

Where GEO actually competes, the answer layer is a long tail, not a cartel. The scary “more extreme than PageRank” headline evaporates the moment you filter to the queries that drive revenue. That’s the first half of the verdict — and it’s good news for anyone who isn’t Reddit.

Then what do the chosen domains actually have in common?

Verdict: not a domain type, not an authority tier — a page format. This was the reversal in Thursday’s results, and it survives scrutiny. Read the corpus one way and it looks vendor-dominated; read it another way and the real mechanism appears:

  • By domain type: tool/vendor own-sites take 59.5% of citations — not the independent review media we predicted.
  • By page format: 66.2% of all URL citations are a ranked best-of listicle or comparison page. In the top 25 most-cited URLs, 92% are.
  • The resolver: vendors win because 61.6% of their citations come from their own best-of / roundup blog content — only 13% from homepages or product pages.

Vendor dominance and format dominance are the same fact seen twice. Engines answering a buyer-intent question reach for the page that already did the ranking work — a ranked, multi-option list — and vendors happen to publish the most of them. The domain type is a red herring. The page format is the signal.

How did the four pre-registered predictions score?

We locked these on Tuesday, before scoring a single domain. Honest tally: one confirmed, three failed — but the one that held is the one that points at the mechanism.

# Pre-registered prediction Result Verdict
1 Review/comparison media + directory/listicle > 50% of citations 25.8% ❌ Failed
2 Top-25 URLs are >60% listicles/comparisons 92% (corpus-wide 66.2%) ✅ Confirmed (strong)
3 ≥3 low-authority niche sites in the top 15 2 ❌ Boundary miss
4 No single domain type exceeds 35% vendor 59.5% ❌ Failed

We bet the anatomy would show up at the domain-type level. It didn’t. It showed up at the page-format level — which we also predicted, and which held at 92%. When a hypothesis is right on one axis and wrong on another, the right axis tells you where the mechanism lives. Prediction #3 is instructive too: structure-beats-size is real, but it lives in the long tail, not the top. Of 219 domains, 170 got a single citation each — no-name niche sites (b2bsaastools, saascrmreview, dataupward) picked purely for publishing a tight best-of page for one specific query. The giants hold the top; well-structured obscure listicles fill everything below.

What should you actually do, per page format?

This is the evergreen payload. Stop optimizing for a domain category and start optimizing for the format engines reach for. Priority order follows the citation share:

Page format Citation share What it means Your move
Best-of listicle / roundup 66% (top-25: 92%) The single citable unit engines reach for Get placed in the roundups that already rank for your category; if none are strong, publish your own tightly-structured “best X for Y”
Comparison (X vs Y) Grouped in the 66% Decision-stage buyer pages Build honest head-to-head pages — including yourself vs. named rivals, with a table
Vendor own best-of blog 62% of vendor citations Vendors win by publishing lists, not product pages Run a content-marketing best-of program on your own domain; rank options honestly
Homepage / product page 8% (13% of vendor) Rarely the cited unit Keep for conversion, not for earning citations — don’t expect engines to quote it
Forum thread (Reddit etc.) 8% Corroboration, not a source See last week’s Reddit verdict — a context signal, not a lever

What makes a best-of page citable — the on-page checklist

  • Ranked, multi-option. A numbered list of real alternatives, not a single-product pitch. Engines cite the page that already compared the field.
  • Query-specific title. “Best CRM for startups,” not “Best CRM.” The long-tail winners matched a narrow buyer query exactly.
  • A comparison table. Structured rows engines can lift into an answer.
  • Recency stamp. “2026” in the title and body — freshness correlates with being chosen (a pattern we keep seeing).
  • Clear per-option structure. One H2/H3 per tool, consistent fields (price, best-for, pros/cons). This is how a low-DR niche page beats a giant in the long tail.

Does this change per engine?

The format signal is consistent, but the volume isn’t. In our corpus Google AI Mode produced the most citations (203) and leans hard on best-of pages plus Reddit presence; Gemini (66) is similarly listicle-heavy; ChatGPT with search on (27) cited fewer sources and left 5 of 12 answers as unlinked brand mentions; Perplexity throttled us to a single answered question, so we under-measure it and say so. The per-format playbook above applies everywhere, but the per-engine nuance — where Reddit helps, where brand mentions travel without a link — is exactly what we mapped in our per-engine GEO checklist and the Reddit verdict. Layer them: format is the shared floor; per-engine tuning is the 30% on top.

What’s the one-line rule to take away?

AI engines cite a page shape, not a domain. For buyer-intent queries, that shape is a ranked best-of list — so get into the ones that already rank, or publish one worth citing. Authority helps you hold the top; structure gets you into the long tail. Neither a big domain nor a niche one is a shortcut — the best-of format is.

What are we testing next?

We now know the format engines choose. Next week we open the listicle itself: what’s inside a citable best-of page? We’ll dissect the on-page anatomy of the pages that got picked — option count, comparison tables, schema markup, recency signals — and test the uncomfortable question this data raised: when vendors rank themselves #1 in their own roundups, do engines catch the self-serving bias, or cite it anyway? Same GEO Lab format: fact-check Monday, pre-registered method Tuesday, results Thursday, verdict Friday.

Frequently asked questions

Do 15 domains really control 68% of AI citations?

For general-knowledge answers, roughly yes — that’s what the 680M-citation industry index measured. For buyer-intent commercial queries, no: our first-party data puts the top 15 at 23%, spread across 220 domains with 77% cited only once. Both are true of different question types; GEO competes in the long-tail one.

If vendor sites take 59.5% of citations, should I optimize my product page?

No — that’s the trap. Only 13% of vendor citations are homepages or product pages. The other 62% are the vendor’s own best-of / roundup blog content. Engines cite the listicle, not the product page. Publish (or get into) ranked, multi-option lists.

What single page format should I prioritize for GEO?

The ranked “best-of” listicle or comparison page. It’s 66% of all citations in our data and 92% of the top 25 URLs. Everything else — homepages, docs, forum threads — is a rounding error by comparison for buyer-intent queries.

I run a small site with low domain authority. Can I still get cited?

Yes, and the long tail proves it: 170 of 219 domains were obscure niche sites cited once each, chosen for publishing a tightly-structured best-of page matching a specific query. Authority holds the top ranks; structure wins the long tail. A precise “best X for Y” list with a comparison table is your path in.

Can I replicate this?

Yes. The 12 buyer-intent questions, three engines, taxonomy rubric, and format audit are all in the pre-registered method post, published before we scored anything, and the raw numbers are in Thursday’s results. Run it on your niche — you’ll likely see the same listicle-dominant format signal.

Sources

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *