What Does a Page AI Cites — But Google Won’t Rank — Look Like? Our Pre-Registered Experiment

Quick answer: This week’s experiment profiles the pages AI engines cite but Google never ranks. From last week’s blind run — 8 buyer-intent queries × 3 engines — we isolated every domain an engine cited that did not appear in that query’s Google top-10. That leaves 45 distinct domains across 62 (query × domain) instances. Before we open a single page, we are pre-registering five falsifiable predictions about what those pages have in common: they are mostly third-party (not the vendor being recommended), they skew toward comparison, listicle, and community formats rather than product homepages, and the mix is more community-and-aggregator-heavy on ChatGPT than on Google AI Mode. If those hold, the “signature” is real and engineerable. If they don’t, we retract — publicly. Data opens Thursday.

Experiment design — published September 8, 2026. This is a pre-registration post in the GEO Lab: we publish the hypothesis, the exact sample, the coding scheme, and the locked predictions now, then open the data later this week. That order is the whole point — you can’t move the goalposts if you plant them in public first.

Why are we running this experiment at all?

Because last week left a loose end. We measured how much AI citations overlap with Google’s top-10, and the answer was: far less than the old “70% overlap” folklore. Blended across three engines the overlap was 34.4%; on ChatGPT it was just 13.3%, while Google AI Mode stayed high at 53.6%. In other words, most of what ChatGPT and Perplexity cite is not what ranks. (Full method and scoring rule are in our pre-registered protocol.)

Measuring the size of a gap tells you it exists. It does not tell you what lives inside it. If a large, non-random share of AI citations comes from pages that don’t rank, the obvious next question is: what do those pages have in common? Are they newer? More structured? Third-party review sites and Reddit threads rather than vendor homepages? If there is a repeatable pattern — a signature — then “get cited without ranking” stops being luck and starts being a spec you can build to. That is the thing a GEO strategy should actually engineer, and it’s exactly what our per-engine checklist is missing a data-backed answer for.

It also matters because the industry just got a shiny new dashboard that can’t answer this. When Google shipped its Generative AI performance report in Search Console, it gave everyone impression counts for Google’s own AI surfaces — but no citation detail, no clicks, and nothing at all about ChatGPT or Perplexity. It tells you a link appeared. It does not tell you what kind of page earns the appearance. To answer that, you still have to instrument it yourself. So we did.

What exactly is the sample — and how do we define “cited but not ranked”?

We reuse last week’s dataset rather than collecting fresh, because the whole value here is profiling the same citations we already scored. For each of the 8 buyer-intent queries and each engine, we recorded two sets: the distinct domains the engine cited, and Google’s top-10 organic domains for that query. A domain is “cited but not ranked” for that query if it appears in the cited set but not in the top-10 set — the set-difference, per our locked overlap rule (Google’s own properties and each engine’s self-citation are excluded, same as last week).

Rolling that up across all queries and engines gives the sample we will code:

Engine Distinct cited-but-not-ranked domains
Perplexity 23
Google AI Mode 18
ChatGPT (web search) 12
Union (distinct) 45
Total (query × domain) instances 62
Cited-but-not-ranked domains from last week’s blind run, 8 buyer-intent queries × 3 engines. A domain can be cited-not-ranked on more than one engine, so per-engine counts sum above the distinct union.

The queries are the same eight commercial shortlisting prompts we’ve used all series: best CRM for small business, best project management software, best email marketing platform, best web hosting for WordPress, best VPN, best password manager, best accounting software for freelancers, and best AI writing tool. These are the queries where “which page gets cited” has real money attached, which is why we keep testing on them instead of trivia.

How will we code each page — and what counts as the “signature”?

Every one of the 45 domains (at the specific cited URL where available) gets coded on five features. We publish the codebook here so the classification isn’t a black box:

  • Party — is the domain the vendor being recommended (first-party, e.g. hubspot.com for “best CRM”) or someone else (third-party: a review site, listicle, forum, news outlet)?
  • Page type — one of: listicle / “best-of”, head-to-head comparison, single review, vendor/product page, editorial or news, community/forum (Reddit, Quora, niche boards), or docs/reference.
  • Structure signals — does the page carry Schema.org markup (Article, Review, ItemList, FAQ, Product) and a scannable, entity-dense list structure, versus prose-only marketing?
  • Freshness — most recent visible publish or update date, where the page exposes one.
  • Brand-name recall — is the domain a recognizable brand/name in its category, or an unknown long-tail site?

Two coders classify independently against this codebook; disagreements are adjudicated and the final codebook plus the full per-domain table ship with Thursday’s results so anyone can re-score us. The “signature” is simply whichever feature values are over-represented among cited-but-not-ranked pages relative to what you’d naively expect. We are not fishing for it after the fact — we are stating below exactly what we expect to find.

What are the five predictions we’re locking before opening the data?

These are pre-registered. Each has a number attached so it can fail cleanly. If the data contradicts them, Thursday’s post says so.

  1. P1 — Third-party dominates. At least 60% of the 45 distinct domains are third-party, not the vendor being recommended. (Rationale: engines cite consensus about a brand, per our finding that brand mentions beat backlinks.)
  2. P2 — List and community formats win. Comparison + listicle + review + community pages together account for at least 55% of the 62 instances, while vendor/product pages are a minority.
  3. P3 — Homepages rarely get cited unranked. A vendor’s own marketing homepage is ≤25% of the cited-but-not-ranked set. If an engine cites a brand’s own site, it usually also ranks; the “gap” is filled by others talking about the brand, not the brand talking about itself.
  4. P4 — The mix is engine-specific. The cited-but-not-ranked set is more community-and-aggregator-heavy on ChatGPT than on Google AI Mode, mirroring the overlap spread (13.3% vs 53.6%). AI Mode’s gap should look more like classic SEO real estate; ChatGPT’s should look more like Reddit-and-reviews.
  5. P5 — Structure over prose. Among pages we can code for it, a majority carry list/Schema structure (ItemList, Review, FAQ, or a clear ranked list) rather than being prose-only marketing copy.

Notice what we are not predicting: we’re deliberately cautious on freshness. We suspect cited-but-not-ranked pages skew newer, but many sites hide or fake dates, so we’ll report freshness descriptively rather than stake a threshold we can’t cleanly verify. Pre-registration means being honest about which claims you can actually settle.

What could make this experiment wrong — and how do we stay honest?

Four limits we’re naming in advance. One: it’s a single logged-out snapshot from last week; AI citations drift, so this is a photograph, not a law. Two: n = 45 domains across 8 categories is enough to see a strong pattern but not to slice finely — we won’t over-claim per-category. Three: ChatGPT contributed the fewest usable queries last week (rate-limiting during collection), so its 12-domain slice is the thinnest and P4 will be read as directional, not precise. Four: page-type coding has judgment in it; that’s why two coders and a published codebook, so you can disagree with our calls line by line.

The reason we design in the open — same as our no-universal-strategy series — is that GEO advice is drowning in confident, untested claims. “Get on listicles.” “Build Reddit presence.” “Add schema.” Maybe. This week we find out whether the pages AI actually cites-without-ranking back those claims or quietly contradict them. Predictions are planted. Thursday we dig.

Frequently asked questions

What does “cited but not ranked” actually mean?

For a given query, it’s a domain that an AI engine cited in its answer but that does not appear in Google’s top-10 organic results for that same query. It’s the set-difference between what AI leans on and what classic search ranks — the part of AI’s answer that traditional SEO position doesn’t explain.

Why pre-register predictions instead of just reporting what you find?

Because open-ended “here’s what the data shows” analysis lets you rationalize any pattern after the fact. By publishing five numbered predictions before we open the pages, we make the experiment falsifiable: if third-party sources aren’t the majority, or homepages dominate, the record already says we were wrong. It’s the same discipline we used to measure the overlap itself.

How big is the sample?

45 distinct domains across 62 (query × domain) cited-but-not-ranked instances, drawn from 8 buyer-intent queries scored against three engines — Perplexity contributed 23 such domains, Google AI Mode 18, and ChatGPT 12. Enough to detect a strong signature; not enough to make fine per-category claims, which we won’t.

Can’t Google’s new AI report in Search Console tell me this?

No. Google’s Generative AI performance report shows impressions for Google’s own AI surfaces only, with no citation detail, no clicks, and nothing about ChatGPT or Perplexity. It tells you a link was shown; it can’t tell you what type of page earns a citation across engines. That gap is exactly why we instrument it by hand.

When do the results come out?

We open and code the data midweek and publish results Thursday with the full per-domain table and charts, then a verdict Friday that judges each of the five predictions and turns any confirmed signature into a per-engine action checklist. This post is the locked design; nothing here changes once results are in.

Sources

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *