We Predicted Test Data Wins the Citation. Cited Sites Scored Lower On All Four Suspects.

🎬 Watch the hands-on test on YouTube, then read the full data and method below.

Quick answer: Our load-bearing prediction (P1) failed. We predicted cited domains would score at least 2× higher than matched comparators on first-party test data. Instead, cited domains averaged 0.45 on a 0–2 scale versus 0.69 for comparators — lower, not higher. That pattern held across all four suspects we scored (test data, schema, brand demand, freshness): the cited group came in below the comparator group on every single one. Two of five pre-registered predictions held (schema barely separates; brand demand is a partial confound), two failed (test data is not the lever we thought; freshness isn’t a clean tie-breaker), and one (P5, cited-but-not-ranked) we couldn’t test with the data we collected. This is the results post for Tuesday’s pre-registered design; nothing below was reframed after the numbers came in.

Experiment results — published September 17, 2026, in the GEO Lab. We locked the query set, the four-suspect rubric, and five falsifiable predictions on Tuesday, collected citations Wednesday, and scored every domain against the rubric before writing a word of this post.

What did we actually score?

We ran 15 buyer-intent “best X” queries — five shopping, five software, five local — logged-out on ChatGPT (web search) and Google AI Mode, the same two engines Tuesday’s design locked. Perplexity sat out this run by design (it rate-limits hard mid-collection, and a half-populated third engine would have muddied the matched comparison). For every query we pulled the distinct domains each engine cited, then paired that set with a matched comparator group — domains that ranked in Google’s organic top 10 for the same query but were not cited by either AI engine. That gave us 98 cited domains and 56 comparator domains (148 unique). We scored every one, 0–2 (schema is 0–1), on the rubric locked Tuesday: first-party test data, schema/structure, brand demand, and freshness. Full scoring criteria are in Tuesday’s post.

Did first-party test data separate cited domains from comparators, like we predicted?

No — and it went the opposite direction from the prediction. P1 said cited domains would score at least 2× higher on test data than their comparators, and be the largest gap of the four suspects. Instead: cited domains averaged 0.45 on test data; comparators averaged 0.69. The comparator group — the organic-ranking sites AI did not cite — scored higher on original testing evidence than the sites it did cite. That is a direct contradiction of the load-bearing prediction, not a narrow miss.

Suspect Cited avg Comparator avg Gap (normalized) Direction
First-party test data (0–2) 0.45 0.69 −0.12 Cited scored lower
Schema/structure (0–1) 0.55 0.67 −0.11 Cited scored lower
Brand demand (0–2) 0.81 1.20 −0.20 Cited scored lower (largest gap)
Freshness (0–2) 1.50 1.74 −0.12 Cited scored lower
Every gap runs the same direction: the comparator group (organic top-10, not cited by AI) outscored the cited group on all four attributes we measured. Gap = (cited avg − comparator avg), normalized to a 0–1 scale so the four suspects are comparable.

The biggest surprise isn’t just that P1 failed — it’s that nothing we measured separates in the predicted direction. If test data, schema, brand demand, or freshness individually explained why AI cites a domain, we’d expect the cited group to lead on at least one. It leads on none. Read the honest interpretation below before concluding “none of it matters” — there’s a real methodology wrinkle in how we built the comparator group that we’re not going to bury.

Did the other three predictions hold up?

P2 — schema barely separates: confirmed. We predicted the schema gap between cited and not-cited domains would be under 15 percentage points. It was 11.3pp. Structure is close to a floor everyone clears, in either direction — consistent with the “AI-ready markup” advice being necessary but not sufficient.

P3 — brand demand is a partial confound: confirmed. We predicted at least 3 cited domains would sit in the low/mid brand-demand tier, and at least 2 high-demand domains would be cited zero times. The actual numbers cleared both thresholds by a wide margin: 68 of the 98 cited domains sit in the low/mid demand tier, and 19 high-demand domains in our comparator pool — including amazon.com, homedepot.com, salesforce.com, and consumerreports.org — were never cited for their query. Brand size alone doesn’t buy a citation; if anything, this run’s biggest gap (brand demand, −0.20) suggests AI is citing smaller, less-mainstream sources than the organic ranking favors, not bigger ones.

P4 — freshness as a tie-breaker inside test data: failed. We predicted that among domains scoring 2 on test data, cited ones would be fresher than not-cited ones, and that no domain scoring 0 on test data and 2 on freshness would get cited. Neither held. Cited domains in the test-data=2 group averaged 1.64 on freshness versus 1.80 for not-cited ones — the opposite of the prediction, though the gap is small. And the “should be zero” case wasn’t close to zero: 33 cited domains scored 0 on test data and 2 on freshness — sites with no original testing evidence that were still cited, apparently on recency alone. You can, in this data, update your way into a citation without anything original to update.

P5 — cited isn’t picked: untested. This prediction needed each engine’s citations tagged by whether the domain was the top recommendation or just a listed source. We collected citation presence, not within-answer ranking, so we can’t score P5 from this run. Flagging it as a genuine gap rather than quietly dropping it — a follow-up collection that captures rank position would close it.

Is this a clean result, or is the comparator group doing something we didn’t intend?

Here’s the limit we want stated plainly, not buried in a footnote. Tuesday’s design called for matching each cited domain to “a like-for-like domain of the same page type that was not cited” — implying deliberate, one-by-one judgment calls (a rival test site, a manufacturer page, a general blog). What we actually built was faster and coarser: for each query, the comparator pool is whatever domains ranked in Google’s organic top 10 but weren’t cited by either AI engine. That’s a reasonable operationalization — those domains genuinely compete for the same query and genuinely weren’t cited — but it isn’t a hand-picked match, and it has a specific bias risk: organic top-10 rankings skew toward high-authority, well-established sites. If AI is citing a different, sometimes smaller-and-newer layer of the web than Google’s organic algorithm rewards, our comparator pool would look bigger-brand and higher-authority than the cited group by construction, independent of whether test data, schema, or freshness are involved at all. The brand-demand gap being the largest of the four (−0.20) is exactly the pattern you’d expect if that’s part of what’s happening.

We’re not walking back the P1 result — a prediction that specific, failing that directionally, is a real finding regardless of comparator methodology, and it directly contradicts Monday’s working theory. But we’re also not going to claim the four-suspect model is exhausted. The honest reading: “first-party test data, alone, does not explain the gap between AI-cited and Google-organic-ranked-but-not-cited domains” is well supported. “Nothing about test data or authority matters to AI citation” is not — that would require a properly hand-matched comparator group we didn’t build this round.

Other limits, named the same way we do every run: this is one logged-out snapshot, on two engines, of 15 queries — not a universal law, and citations drift. Scoring is a judgment call on a published rubric, done in parallel by independent scorers per domain, not by a single grader; we’re releasing the full 148-domain score sheet so anyone can re-score. Structured-data detection leaned on automated fetches that sometimes couldn’t see server-rendered JSON-LD, so schema scores are the least confident of the four (flagged 0 in several cases where a manual browser check might find valid markup) — treat the 11.3pp schema gap as directionally right, not to the decimal.

So what actually explains the citations, if not these four?

We don’t know yet, and we’d rather say that than force a fifth suspect into the data after the fact — that’s the exact move Tuesday’s post promised not to make. What we can say: the sites AI cited leaned toward independent review/testing outlets (RTINGS, Vacuum Wars, HouseFresh, SoundGuys, Security.org), niche buyer’s-guide sites, and — in local queries — a long tail of small individual business pages that showed up for none of our four suspects in a way that separated them from comparators. That’s consistent with something we keep bumping into across this series: there is no single universal GEO lever, and query-category-specific dynamics (shopping vs. software vs. local) may matter more than any single cross-category attribute. Friday’s verdict will lay out what we’d test next given this result, rather than declare a winner the data doesn’t support.

Frequently asked questions

Does first-party test data explain why AI cites certain domains?

Not on its own, based on this test. We predicted cited domains would score at least 2× higher than matched comparators on first-party test data; instead cited domains averaged 0.45 versus 0.69 for comparators on a 0–2 scale — lower, not higher. The prediction failed in the opposite direction from what we expected.

Did any of the five pre-registered predictions hold?

Two held cleanly: schema barely separates cited from not-cited domains (an 11.3 percentage-point gap, under the 15pp threshold we set), and brand demand is a confound rather than the lever (68 of 98 cited domains sit in the low/mid demand tier, and 19 high-demand domains were never cited). Two failed: test data (P1, reversed direction) and freshness-as-tie-breaker (P4, both sub-clauses contradicted). One (P5, cited-vs-ranked) couldn’t be tested with this run’s data.

Why did cited domains score lower than comparators on every attribute?

We don’t have a confirmed explanation, and we’re naming the most likely confound rather than guessing past it: our comparator group was built from Google’s organic top-10 rankings for each query, which skew toward high-authority, well-established sites. If AI cites a somewhat different, sometimes smaller layer of the web than the organic algorithm rewards, the comparator pool would look higher-scoring on every attribute by construction — independent of what actually drives citation. The result that first-party test data specifically isn’t the lever still stands; a claim that “nothing about authority matters” would need a hand-matched comparator group we didn’t build this round.

Why is Perplexity missing from this experiment?

It was excluded by design before collection started, not dropped afterward. Perplexity rate-limits anonymous queries hard mid-collection, and a half-populated third engine would have added noise to the matched comparison rather than signal. This is a two-engine run: ChatGPT web search and Google AI Mode.

What happens if a locked prediction fails — do you rewrite it?

No. The rubric and predictions were published Tuesday, before any domain was scored, specifically so they couldn’t be adjusted after seeing the data. When P1 came back reversed, we reported the reversal rather than reframe the suspect list or redefine the scoring bands. That is the entire discipline of pre-registration.

Sources

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

🦜 Follow GeoParrot: YouTubeX