Quick answer: A widely shared new index — 5WPR’s AI Platform Citation Source Index 2026, built from 680 million citations across six studies — claims the top 15 domains capture 68% of all AI citation share, calling the concentration “more extreme than Google PageRank ever produced.” We reran the exact same concentration math on our own first-party data from buyer-intent queries. The result is a different world: the top 15 domains capture just 23%, citations spread across 220 distinct domains, and 77% of those domains were cited only once. Both numbers are real — they just measure different questions. The 68% describes general-knowledge AI answers dominated by Reddit and Wikipedia. The queries your business actually competes for are a long tail. This week we dissect what that long tail is made of: the anatomy of the domains AI engines actually choose.
This is Monday’s fact-check, and it opens a new week-long question in the GEO Lab: which domains do AI engines actually choose to cite for commercial queries — and is the “a handful of domains control everything” story true where it matters? We’ll spend the week taking the chosen citation apart.
What does the industry index actually claim?
The headline making the rounds this month comes from 5WPR’s index, which synthesizes 680 million individual citations across ChatGPT, Google AI Overviews, Perplexity, Gemini, and Claude, drawn from six studies run between August 2024 and April 2026. Its punchiest findings:
- Top 15 domains capture 68% of all consolidated AI citation share.
- Reddit sits at roughly 40% frequency across LLMs.
- Wikipedia accounts for 26–48% of ChatGPT’s top-10 citation share.
- Concentration is “far more extreme than Google PageRank ever produced.”
Taken at face value, that’s a scary picture for anyone doing generative engine optimization: if 15 sites own two-thirds of the answer layer, the game looks rigged before you start. But there’s a quiet word doing enormous work in that sentence — consolidated. When you consolidate 680M citations across every kind of query — “who won the 1998 World Cup,” “is coffee bad for you,” “what is photosynthesis” — you’re mostly measuring general knowledge. And general knowledge is concentrated: it lives on Reddit, Wikipedia, and major news. That’s a real finding. It’s just not the finding a business needs.
What did our own buyer-intent data show?
Businesses don’t get bought through “what is photosynthesis.” They get found through commercial, buyer-intent queries — “best GEO tools,” “X vs Y,” “top AI visibility platforms.” So we asked the concentration question of those queries specifically. Using the raw citations from our provenance experiment (three AI engines × 12 buyer-intent questions, every source URL preserved), we ran the identical top-15 math:
| Concentration metric | Industry index (all queries) | Our data (buyer-intent) |
|---|---|---|
| Top 15 domains’ share | 68% | 23% |
| Reddit’s share | ~40% frequency | 3.8% |
| Distinct domains cited | concentrated | 220 |
| Domains cited only once | — | 170 of 220 (77%) |
| Single most-cited domain | Reddit — but only 3.8% |
Same math, one-third the concentration. For the queries GEO actually targets, no domain owns the answer layer — not even the #1. The citation basket is a long tail: 220 different domains for just 12 questions, and more than three-quarters of them showed up exactly once. This is the same fragmentation we found when we measured cross-engine overlap — where only 0.5% of sources were shared across four engines. Concentration and overlap are two different lenses, and both point the same way: for commercial queries, the citation layer is diffuse.
Why is there a 3x gap between the two numbers?
The index isn’t wrong and our data isn’t wrong — they answer different questions. Four things drive the gap:
- Query mix. Aggregate indexes blend informational queries (where Reddit/Wikipedia genuinely dominate) with commercial ones. Filter to buyer-intent and the concentration collapses, because there’s no single “Wikipedia of best-CRM-software.”
- Frequency vs. share. “Reddit appears at 40% frequency” counts how often Reddit shows up at all. “Reddit is 3.8% of sources” counts its slice of every citation. In our data Reddit was the most common single domain yet a tiny share — because engines cite ~17 domains per answer. Presence isn’t dominance.
- Time volatility. The index itself notes ChatGPT’s Reddit citations “fell from roughly 60% to 10% in six weeks” in late 2025. A number aggregated over 20 months smooths away swings that would blow up any fixed strategy. As the report puts it, volatility happens “within weeks, not years.”
- It’s not a moat you can buy. In May 2026 Google published its own guidance stating that optimizing for generative AI is “still SEO,” and that manufactured mentions and llms.txt files “aren’t as helpful as they might seem.” Concentration in the aggregate doesn’t mean you win by chasing the 15 “power domains” — it means you earn citations the same way you always earned rankings: by being the best answer.
So which domains DO get chosen for buyer queries?
Here’s the part that sets up the rest of the week. When we look at what the engines actually reached for on commercial questions, the top of our list isn’t Reddit and Wikipedia. It’s category media and tool directories:
- TechRadar (10), Zapier (6), PCMag (6) — commercial-review media
- AllAboutCookies (5), ZDNet (4), Security.org (3) — niche authority/comparison sites
- Forbes (3), LinkedIn (4) — brand/business surfaces
- …and a long tail of category-specific SaaS-comparison and listicle domains, each cited once or twice
That’s a completely different citation economy from the general-knowledge one. It’s not “own Wikipedia.” It’s “be the kind of source a review roundup, a comparison page, or a category authority would cite — or be that page.” That’s the thread we pull all week: what do the domains AI engines actually choose have in common? Tuesday we pre-register the method for dissecting them; Thursday we publish the anatomy; Friday, the verdict and a per-type playbook.
What should this change about your GEO strategy?
Three practical shifts if you’ve been reading the “68% concentration” headline as marching orders:
- Don’t chase the 15 “power domains.” For your buyer’s queries, they don’t dominate. Getting a Reddit mention or a Wikipedia edit is not a citation strategy — our Reddit verdict showed a mention is a contributing signal, not a citation lever.
- Win the long tail. 77% of cited domains appeared once. That fragmentation is opportunity: category-relevant media, comparison pages, and well-structured pages of your own can all break in, because there’s no incumbent to unseat.
- Get cited on the chosen domains, and become one. The realistic play is earning mentions inside the review media and comparison directories engines already trust for your category — while making your own pages clean enough to be cited directly.
FAQ
Is the 68% figure wrong, then?
No. For an aggregate of 680M citations spanning all query types, it’s plausible and useful — general-knowledge answers really do concentrate on Reddit, Wikipedia, and major news. Our point is narrower: it doesn’t describe the buyer-intent queries a business competes for, where our data shows ~23%.
Should I try to get on Reddit or Wikipedia for GEO?
Only if it’s genuinely warranted. In our buyer-intent data Reddit was 3.8% of sources and Wikipedia barely registered. Community presence can help indirectly, but neither is a reliable citation source for commercial queries. Chasing them as a shortcut misreads the aggregate number.
Which domains should I target instead?
The ones engines already choose for your category — review media (TechRadar, PCMag, ZDNet-style), comparison and directory pages, and category-specific authorities — plus your own well-structured pages. This week’s experiment breaks down what those chosen domains share.
How did you measure the 23%?
We took every source URL three AI engines cited across 12 buyer-intent questions (312 citations, 220 distinct domains after excluding an engine’s own navigational self-links), summed the top 15 domains’ counts, and divided by the total — the same concentration calculation aggregate indexes use. It’s a small, hand-run sample by design; treat it as “the pattern is strong,” not “23% to the decimal.” The full method and data are published.
Does concentration change over time?
Yes — sharply. Even the industry index flags that ChatGPT’s Reddit share fell from ~60% to ~10% in six weeks. Any strategy built on “these 15 domains” is betting on a snapshot that moves within weeks.
Sources
- 5WPR — AI Platform Citation Source Index 2026 (680M citations; the 68% / 15-domain claim)
- GEO Lab — Reddit Is the #1 Domain AI Engines Cite — and Still Only 4% of Their Sources (this post’s first-party data)
- GEO Lab — We Asked 4 AI Engines the Same 15 Questions — Only 0.5% of Sources Overlapped (cross-engine fragmentation)
- GEO Lab — Does a Reddit Mention Cause an AI Citation? The Verdict (why “get on Reddit” oversells)
- GEO Lab — We Asked 3 AI Engines the Same 12 Questions (our first citation run)

Leave a Reply