Quick answer: We took the 306 citations across 219 domains three AI engines gave for 12 buyer-intent questions and dissected what the chosen domains have in common. The result flips depending on how you look. By domain type, tool/vendor own-sites dominate — 59.5% of all citations — not the independent review media we predicted. But by page format, 66% of every citation is a “best-of” listicle or comparison page, and even vendors earn 61.6% of their citations from their own roundup blog content, versus just 13% from homepages or product pages. So the unit AI engines actually reach for isn’t a kind of site — it’s a kind of page: a ranked, multi-option list. Two of our four pre-registered predictions failed on the type axis and one held strongly on the format axis, and that split is the whole story.
On Monday we showed 15 domains don’t control AI citations for buyer-intent queries — the top 15 capture 23%, spread across a 220-domain long tail. On Tuesday we pre-registered a three-part method and locked four predictions before scoring a single domain. This post is the raw result, graded against what we said would happen. As promised, we re-analyzed the exact corpus from our provenance run — no fresh pull — so the numbers stay comparable to Monday’s 23%.
What kind of domain do engines actually cite?
Part 1 sorted all 219 cited domains into six buckets, fixed before scoring, and measured each bucket two ways: its share of the 306 citations, and its share of distinct domains. Our prediction — that independent review/comparison media plus directory/listicle sites would together capture over half the citations — failed hard:

- Tool / vendor own-sites: 59.5% of all citations (152 of 219 domains). The single biggest bucket by far.
- Directory / listicle: 13.1% — the “best 10 X” roundups and SaaS aggregators.
- Review / comparison media: 12.7% — TechRadar, PCMag, ZDNet and friends. We predicted this would lead; it came third.
- Community/UGC 6.5%, niche authority 6.5%, publisher 1.6%.
Review media + directory together are only 25.8% — nowhere near the >50% we called (Prediction #1: ❌). And no “single type stays under 35%” ceiling survived either: vendor own-sites alone take 59.5% (Prediction #4: ❌). Read at the domain-type level, the answer layer looks vendor-dominated. Which sounds like the opposite of a useful GEO lesson — until you change the axis.
What kind of page do engines actually cite?
Part 2 ignored who owns the page and asked what the page is. We audited the format of every cited URL — best-of listicle, comparison, homepage/product page, doc, forum thread, or other article. This is where the pattern snaps into focus:

Across 396 URL-level citations, 66.2% are “best-of” listicles or comparison pages — a ranked, multi-option page that answers “what are the best tools for X.” Homepages and product pages? 8%. Forum threads (Reddit et al.)? 8%. That’s Prediction #2, and it held strongly: in the top 25 most-cited URLs, 92% were listicles or comparisons (✅, well over the 60% we called). Engines answering a buyer-intent question overwhelmingly reach for the page that already did the ranking work for them.
This is the resolution to the type-axis confusion. Vendors dominate not because engines love product pages — but because vendors are the ones cranking out the most “best-of” content marketing.
Why do vendors win — product pages, or listicles?
If vendor own-sites take 59.5% of citations, you’d assume engines are citing homepages and product pages. They aren’t. When we split vendor citations by page format, the reversal is unmistakable:

61.6% of vendor citations come from the vendor’s own best-of or roundup blog content — a SaaS company publishing “the 10 best CRMs for startups” and landing in the answer. Only 13% are homepages or product pages. So the vendor dominance and the format dominance are the same phenomenon seen twice: engines choose ranked, multi-option pages, and vendors happen to publish the most of them (often ranking themselves #1, but that’s a different post). The domain type is a red herring; the page format is the signal.
Does high domain authority decide who gets chosen?
Part 3 asked whether being chosen requires raw domain authority. In the top 15 most-cited domains, we predicted at least three would be low-authority niche sites — structure beating size. We found two (dageno.ai, usebear.ai), just under the line (Prediction #3: ❌, boundary). But the honest read isn’t “authority wins.” It’s that structure-beats-size shows up in the long tail, not the top: of the 219 domains, 170 (77.6%) were cited exactly once, and that long tail is packed with no-name niche sites — b2bsaastools, saascrmreview, dataupward, wdmarket — that got picked purely because they published a tightly-structured best-of page for a specific query. The giants hold the top; obscure, well-structured listicles fill everything below it. That’s the same 23% concentration / 220-domain sprawl we found Monday, now explained.
How did the four predictions score?
| # | Pre-registered prediction | Result | Verdict |
|---|---|---|---|
| 1 | Review/comparison media + directory/listicle > 50% of citations | 25.8% | ❌ Failed |
| 2 | Top-25 URLs are >60% listicles/comparisons | 92% (corpus-wide 66.2%) | ✅ Confirmed (strong) |
| 3 | ≥3 low-authority niche sites in the top 15 | 2 | ❌ Boundary miss |
| 4 | No single domain type exceeds 35% | vendor 59.5% | ❌ Failed |
Three of four “failed” — and that’s the finding, not a defeat. We predicted the anatomy would show up at the domain-type level (review media wins). It didn’t. It showed up at the page-format level (listicles win) — which we also predicted, and which held at 92%. When a hypothesis is right on one axis and wrong on another, the axis that’s right tells you where the real mechanism lives. Here: engines select page format, not site type.
What’s the honest takeaway before Friday?
One clean result: for buyer-intent queries, the domains AI engines choose share a page format — a ranked, multi-option “best-of” list — far more than they share a domain type. Vendor sites lead the citation count, but 62% of the way they earn it is by publishing that same listicle format, not by fielding product pages. This lines up with last week’s “presence ≠ source” finding and the fragmentation across engines: there’s no shortcut through authority or site category — there’s a page shape engines reach for. On Friday we turn this into a verdict and a per-format playbook: get into the best-of lists that already rank, or publish one worth citing.
Two caveats we won’t bury. Perplexity throttled us — it gave usable sources for only 1 of 12 questions — so the corpus leans Google/Gemini/ChatGPT, and we say so rather than average it away. And this is a small, hand-run sample: treat every number as “the pattern is strong,” not “these percentages to the decimal.” The full 219-domain label set is reproducible from the method post.
FAQ
Vendors dominate at 59.5% — so should I just optimize my product page?
No — that’s the trap this data springs. Only 13% of vendor citations are homepages or product pages. The other 62% are the vendor’s own best-of / roundup blog content. Engines are citing the listicle, not the product page. Publish (or get into) ranked, multi-option lists.
How can review media “lose” if TechRadar and PCMag are top domains?
Both are true. Review media is only 12.7% of citations by volume, but individual review sites (TechRadar, PCMag, ZDNet) rank high as single domains because each publishes many strong best-of pages. Type-share and per-domain rank measure different things — the same “presence ≠ dominance” gap we keep hitting.
Why did three of four predictions fail?
Because we bet the pattern would appear at the domain-type level (review media leads) and it appeared at the page-format level (listicles lead) instead. Pre-registration means we report that honestly. The prediction that held — 92% of top URLs are listicles — is the one that points at the real mechanism.
Isn’t 306 citations too small to trust?
It’s a deliberately small, transparent sample — three engines × 12 buyer-intent questions, every URL preserved. Perplexity throttled to 1 of 12, so it’s under-represented. Read it as a strong directional pattern, not decimal-precise truth, and replicate it on your own niche.
Can I replicate this?
Yes. The 12 questions, three engines, taxonomy rubric, and format audit are all in the method post, published before we scored anything. Run it on your category — you’ll likely see the same listicle-dominant format signal.
Sources
- GEO Lab — What Do the Domains AI Engines Choose Have in Common? Our Pre-Registered Method (this experiment’s pre-registration)
- GEO Lab — Do 15 Domains Control 68% of AI Citations? Our Buyer-Intent Data Says 23% (this week’s claim)
- GEO Lab — Reddit Is the #1 Domain AI Engines Cite — and Still Only 4% of Their Sources (the source corpus)
- GEO Lab — We Asked 4 AI Engines the Same 15 Questions — Only 0.5% of Sources Overlapped (citation fragmentation)
- GEO Lab — Does AI Search Use Backlinks? (authority vs structure)

Leave a Reply