What Do the Domains AI Engines Choose Have in Common? Our Pre-Registered Method to Dissect Them

Quick answer: Monday we found that for buyer-intent queries, no small set of domains “controls” AI citations — the top 15 capture just 23%, spread across 220 distinct domains. So the useful question isn’t who dominates; it’s what the domains engines actually reach for have in common. This week’s experiment dissects them. Our pre-registered method has three parts run on the same 220-domain first-party corpus: (1) a domain-type taxonomy — bucket every cited domain into review media, tool/vendor site, directory/listicle, community/UGC, publisher, or niche authority; (2) a page-format audit — is the cited page a “best-of” listicle, a comparison table, a homepage, a doc, or a forum thread?; and (3) an authority-vs-structure test — do chosen domains win on domain authority, or can fresh, tightly-structured niche pages break in? We’ve locked four predictions before scoring a single domain. This post is the method, published before the results, so you can hold us to it.

Yesterday’s fact-check opened this week’s question: an industry index claims 15 domains capture 68% of AI citations, but our buyer-intent data says 23% — a 220-domain long tail where 77% of domains were cited exactly once. That’s reassuring if you’re a small brand fighting for visibility, but it raises the harder question: if the answer layer isn’t owned by a handful of giants, then which domains do engines choose for commercial queries, and is there a pattern you can actually build toward? This week we take the chosen citation apart. Today, before we touch the scoring, here’s exactly how — with the same pre-registration discipline we used for our Reddit causation method and our cross-engine overlap method.

What question are we actually answering?

One testable claim, stated plainly:

Hypothesis (H1): AI engines don’t reach for cited domains at random. For buyer-intent queries they disproportionately choose a specific kind of page — independent review/comparison media and structured “best-of” listicles — over brand homepages, docs, or general-knowledge sources. The domains engines choose share a common anatomy: category-relevant, multi-option, structured, and recent.
Null hypothesis (H0): Chosen domains are just a reflection of raw domain authority — the biggest, highest-backlink sites win, format and type don’t matter, and there’s no repeatable pattern a smaller brand could engineer toward.

If H1 holds, GEO has a concrete target: be the type of page engines choose, or get cited inside it. If H0 holds, the honest advice collapses back to “just build authority the slow way.” The method below is built to tell those two worlds apart.

What corpus are we dissecting — and why reuse it?

We’re not collecting a fresh pull for this one. We’re re-analyzing the exact citation corpus from our provenance experiment: three AI engines × 12 buyer-intent questions, every source URL preserved — 410 citations across 220 distinct domains. Reusing it is deliberate: the concentration number Monday (top 15 = 23%) came from this same corpus, so the anatomy has to be measured on the same data or the two posts wouldn’t line up. Collecting new answers would introduce a different query mix and make the week’s numbers incomparable. The tradeoff, stated up front: it’s a small, hand-run sample — treat every result as “the pattern is strong,” not “these percentages to the decimal.” Where an engine throttled us (Perplexity gave usable sources for only 1 of 12 questions last run), that engine is under-represented and we’ll say so, never hide it.

What are the three parts of the dissection?

Part What it measures What a “there’s a common anatomy” result looks like
1 · Domain-type taxonomy What kind of site each cited domain is A few types (review media + directory/listicle) capture the majority of citations
2 · Page-format audit What kind of page each cited URL is “Best-of” listicles and comparison pages dominate over homepages and docs
3 · Authority-vs-structure Whether high domain authority is required to be chosen Fresh, low-authority niche pages appear alongside the giants — structure beats size

Part 1 — The domain-type taxonomy

Every one of the 220 cited domains gets sorted into exactly one bucket, using a rubric fixed before scoring:

  • Review / comparison media — independent editorial that rates and ranks tools (TechRadar, PCMag, ZDNet).
  • Tool / vendor own-site — the product’s own domain (e.g. a SaaS company’s homepage or blog).
  • Directory / listicle / aggregator — “best 10 X” roundups and category directories, often from smaller SEO-driven sites (b2bsaastools, saashero-style).
  • Community / UGC — forums and social (Reddit, LinkedIn, YouTube).
  • Publisher / news — general business/tech press (Forbes).
  • Niche authority — specialist single-topic sites (Security.org, AllAboutCookies).

We report each type’s share of total citations and its share of distinct domains — because a type can be common (many domains) yet low-share (each cited once), or rare yet heavily cited. That gap is exactly the “presence ≠ dominance” trap we hit with Reddit’s 4% source share, and we’re measuring both sides on purpose.

Part 2 — The page-format audit

Type isn’t format. TechRadar can publish a listicle or a single-product review; a vendor site can be a homepage or a comparison page. So for the top ~25 most-cited URLs we hand-code the page format: “best-of” listicle, head-to-head comparison, single-product page/review, documentation/guide, homepage, or forum thread. The hypothesis is that engines reach for multi-option, extractable pages — the ones that already contain a ranked answer — far more than single-brand pages. If most cited pages are listicles and comparisons, the GEO play is obvious; if they’re homepages, it’s a different game entirely.

Part 3 — The authority-vs-structure test

This is the part that decides between H1 and H0. For the top 15 most-cited domains we record two things: a public domain-authority bucket (high / medium / low, from a standard third-party metric) and whether the cited page carries structure + recency signals — a visible 2025–2026 publish/update date, H2 question headers, a comparison table, a numbered list. If the top-cited set is uniformly high-authority, the null wins: you need to be big. If low-authority niche sites (young domains, thin backlink profiles) sit right next to the giants because their pages are fresh and tightly structured, that’s direct evidence structure can substitute for size — the most actionable finding this week could produce, and consistent with what we saw when backlinks turned out to be a weaker AI-citation signal than in classic SEO.

How do we score without fooling ourselves?

Risk How we control for it
Motivated classification The type + format rubric is written and frozen before scoring; predictions locked below
Ambiguous domains Every domain gets one primary bucket; edge cases are logged with the reason, not silently forced
Coder drift Two passes over the same list; disagreements re-checked against the rubric, not re-argued
Long-tail noise Type taxonomy runs on all 220 domains; format + authority audits focus on the repeatedly-cited top, where a pattern would actually matter
Engine imbalance Per-engine coverage reported (Perplexity throttled to ~1/12 last run); we never average away a gap

The raw domain list, buckets, and scores will ship as a downloadable table with the results, so anyone can re-run the counts — the same reproducibility standard as our cross-engine data.

Our predictions — locked in before we score a single domain

Pre-registration is what separates an experiment from a story told afterward. On the record, before scoring:

  1. Review/comparison media + directory/listicle types will together capture over 50% of citations — more than brand-own-sites and community/UGC combined. Engines reach for third-party roundups, not homepages.
  2. Among the top ~25 cited URLs, over 60% will be “best-of” listicles or head-to-head comparison pages — multi-option, ranked, extractable — not single-product pages or homepages.
  3. Structure beats size: at least 3 of the top 15 most-cited domains will be low-authority niche/directory sites (young, thin backlink profiles), proving fresh well-structured category pages break in next to the giants.
  4. No single domain type will exceed ~35% of citations. The chosen set is a mix, echoing Monday’s long tail — there’s a common anatomy (type + format + recency) but not a single dominant kind of site.

If the data contradicts any of these, we’ll say so on Friday. That’s the entire point of writing them down where you can see them.

Next at GEO Lab: we score all 220 domains against this rubric midweek and publish the anatomy Thursday — every number against these four predictions, with charts — then deliver Friday’s verdict and a per-type playbook: given what engines actually choose, where should a brand spend its GEO effort? Start with Monday’s fact-check →

FAQ

What are you actually testing this week?
Not whether citations are concentrated — Monday answered that (they’re not, for buyer-intent queries). This week tests what the cited domains have in common: their type, their page format, and whether raw domain authority or page structure decides which ones engines choose.

Why reuse old data instead of collecting fresh answers?
Because Monday’s concentration number and this week’s anatomy have to come from the same corpus, or the numbers wouldn’t be comparable. A fresh pull would change the query mix. We reuse the 220-domain provenance corpus and treat it as a strong-signal sample, not a precise census.

Isn’t “domain type” subjective?
Somewhat — which is exactly why the rubric is frozen before scoring, every domain gets one primary bucket, ambiguous cases are logged, and the full labeled list ships with the results so you can disagree with any specific call.

What would prove the null hypothesis?
If the top-cited domains are uniformly high-authority and page format doesn’t matter — i.e., engines just cite the biggest sites — then there’s no engineerable “anatomy,” and the honest advice is simply to build authority. We’ve predicted the opposite; Friday will show which way it broke.

When do we get the answer?
The anatomy results and charts land Thursday, with the verdict and per-type playbook Friday.

Sources

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *