Is It the Test Data, the Schema, the Brand, or the Freshness? Our Pre-Registered Test of Why AI Defaults to Independent Test Labs

Quick answer: Monday’s fact-check made a claim we like a little too much — that AI defaults to RTINGS and TechRadar because of first-party test data, not schema. Liking a hypothesis is exactly when you should try to break it. So this week we run four suspects in a lineup: first-party test data, schema/structure, brand demand, and freshness. Across 15 buyer-intent queries on two engines, we’ll score every domain the engines cite — plus matched comparators they don’t cite — on all four attributes using a rubric we publish below, then check which one actually separates cited from not-cited. We think test data wins by a mile and schema barely moves the needle. Five numbered, falsifiable predictions are locked before we touch the data. Collection starts Wednesday.

Experiment design — published September 15, 2026, in the GEO Lab. This is a pre-registration: the hypothesis, the query set, the scoring rubric, and five locked predictions go out now; we collect and open the data later this week. Planting the goalposts in public before kickoff is the whole discipline — and the only honest way to test a hypothesis we’re rooting for.

What claim are we trying to break this week?

Monday’s take was clean: for “best X” queries, AI reaches for independent test labs — RTINGS, TechRadar, PCMag — plus YouTube and Reddit, and it does so because those sites own first-party test data (measured input lag, real battery drain, hands-on footage) that no amount of schema markup can manufacture. Schema is the floor everyone clears; original measurement is the moat.

Neat story. But it’s observational, and there are at least three rival explanations that would produce the exact same citation list. Maybe AI cites RTINGS because it’s a big, high-demand brand the model has seen a million times — not because of the data inside. Maybe it cites TechRadar because those pages are freshly updated and the manufacturer’s page is stale. Maybe the labs just happen to have cleaner structure and schema than the pages that lose. If any of those is the real driver, “go publish original test data” is the wrong advice — and we’d be guilty of the same confident, untested storytelling we criticize. So we put all four suspects in the room and see which one has an alibi.

Who are the four suspects — and how do we score each one?

The design is a matched comparison. For each query, we pull the distinct domains each engine cites (the “cited” group), then pair each with a like-for-like domain of the same page type that was not cited for that query (the “comparator” group) — a rival test site, a manufacturer product page, a general blog. Then we score every domain in both groups on four attributes, using a rubric fixed in advance so scoring can’t drift toward the answer we want:

Suspect Scoring rubric (locked) What it would mean if it’s the lever
1. First-party test data 0 = restates specs, no original testing · 1 = first-hand/hands-on experience (used it, subjective) · 2 = standardized lab measurements published (numbers nobody else has) Cited domains carry original data; you win by measuring things
2. Schema / structure 0 = no structured data · 1 = valid Product/Review/Article schema + parseable spec block (checked with a structured-data validator) You win by marking up your page — the “AI-ready” advice is right
3. Brand demand 0 = low branded-search tier · 1 = mid · 2 = high (bucketed from a public authority/traffic estimate, disclosed per domain) AI cites who it already knows; small sites can’t break in
4. Freshness 0 = last meaningful update >18 months · 1 = 6–18 months · 2 = <6 months You win by updating constantly, regardless of what’s on the page
Every cited and comparator domain is scored 0–2 (schema 0–1) on all four suspects, two-pass, before any cross-tabulation. The suspect with the biggest gap between the cited and not-cited groups is the strongest candidate lever.

The logic is simple: if first-party test data is really the lever, cited domains should score high on suspect 1 while their non-cited comparators score low — a wide gap. If schema were the lever, we’d see the gap on suspect 2 instead. Because most competitive pages already clear the eligibility floor, we expect suspect 2 to show almost no gap — cited and not-cited pages both have schema. That’s the whole test in one sentence: which column separates the winners from the losers?

What exactly gets collected, and on which engines?

Fifteen buyer-intent queries, three buckets of five, chosen to match where product recommendations actually happen:

  • Shopping (5): best robot vacuum, best OLED TV under $1500, best noise-cancelling headphones, best office chair, best air purifier.
  • Software (5): best CRM for small business, best project management tool, best password manager, best email marketing platform, best VPN.
  • Local services (5): best dentist in Seattle, best plumber in Denver, best gym in Chicago, best coffee shop in Austin, best moving company in Phoenix.

We run all 15 on ChatGPT (web search) and Google AI Mode — the same two engines Monday’s snapshot used, so this extends that dataset rather than starting a new one. Collection is logged-out, a single snapshot, excluding each engine’s self-citations and Google’s own properties (the exclusions we’ve used all series). Perplexity sits out this run: it rate-limits hard mid-collection, and a half-populated third engine would muddy the matched comparison more than it helps. The full per-domain score sheet ships with Thursday’s results.

What five predictions are we locking before collection?

Pre-registered, numbered, each falsifiable. If the data contradicts one, Thursday’s post says so in plain language — no quiet reframing.

  1. P1 — Test data separates hardest. The mean first-party-data score (suspect 1) of cited domains is at least 2× that of their matched non-cited comparators, and this is the largest cited-vs-not gap of the four suspects. This is the load-bearing prediction: if it fails, Monday’s thesis fails.
  2. P2 — Schema barely separates. The share of domains carrying valid schema differs by less than 15 percentage points between the cited and not-cited groups (both high). Structure is a floor everyone clears, not the thing that breaks the tie.
  3. P3 — Brand demand is a confound, not the lever. At least 3 cited domains sit in the low or mid demand tier, and at least 2 high-demand domains (big manufacturer or retailer pages) are cited zero times. Brand size alone leaves the split unexplained.
  4. P4 — Freshness is a tie-breaker inside the lever, not a substitute for it. Among domains scoring 2 on test data, cited ones are fresher than not-cited ones — but no domain scoring 0 on test data and 2 on freshness gets cited. You can’t update your way into a citation with nothing original to update.
  5. P5 — Cited isn’t picked. At least 30% of cited domains appear as sources but are not the engine’s top recommendation — the cited-but-not-ranked layer — and lab-data domains skew toward “picked” while forums and YouTube skew toward “cited-but-not-ranked.” Being a source and being the answer are different rungs.

Together these describe a specific shape: first-party data is the lever, schema is the floor, brand demand is a partial confound, and freshness only matters once you already hold original data. That’s the opposite of the “add schema and a review widget” narrative. It also lines up with the 2026 ranking-factors survey crowning original research the top content sub-factor, and with why earned reputation beats mechanical signals. If P1 fails — if schema separates as well as test data, or a fresh spec-restater gets cited — we were wrong, and Friday’s verdict will say it out loud.

What could make this experiment misleading — and how do we stay honest?

Four limits, named in advance. One: it’s observational, not causal. We can’t randomly assign a site “more test data” and watch citations move; we can only show which attribute correlates with citation. A wide gap on suspect 1 is strong circumstantial evidence, not proof — we’ll frame it that way. Two: the rubric involves judgment. “Standardized lab measurements” versus “hands-on” is a call, so we publish the rubric, score every domain twice, and release the full sheet for anyone to re-score. Three: the brand-demand proxy is imperfect — a public authority estimate is a stand-in for true search demand, and we’ll flag any domain where it’s borderline. Four: it’s one logged-out snapshot of 15 queries on two engines — a strong signal in one place and time, not a universal law, and AI citations drift and vary by engine.

We design in the open — same as the rest of the GEO Lab — precisely because “publish original test data” is our own advice, and advice you’re attached to is the easiest to fool yourself about. The matched comparison and the locked predictions are there to make it possible for the data to embarrass us. Predictions are set. Wednesday we collect; Thursday we open the score sheet; Friday we judge all five against what we said here.

Frequently asked questions

What is this experiment actually testing?

Which of four attributes — first-party test data, schema/structure, brand demand, or freshness — best separates the domains AI cites for “best X” queries from comparable domains it doesn’t cite. It’s the measured follow-up to Monday’s claim that original test data, not markup, is why AI defaults to labs like RTINGS and TechRadar.

Why score non-cited “comparator” domains too, instead of just the cited ones?

Because a list of cited domains alone can’t tell you what caused the citation — cited sites might have schema, but so might everyone. Pairing each cited domain with a like-for-like page that wasn’t cited isolates the attribute that differs. If schema is present in both groups equally, schema isn’t the lever; if test data is high only in the cited group, it’s the prime suspect.

Why pre-register the predictions instead of just reporting results?

Open-ended “here’s what we found” analysis lets you rationalize any pattern after the fact — especially when you’re rooting for a hypothesis. Publishing five numbered predictions before collecting makes the test falsifiable: if schema separates as strongly as test data, or a fresh spec-restater gets cited, the record already says we were wrong. It’s the discipline we use for every GEO Lab run.

Can a correlation study really prove first-party data causes citations?

No, and we don’t claim it can. This is observational: we can show which attribute travels with citations, not prove it causes them, because we can’t randomly assign test data to a site. A wide, consistent gap on the test-data suspect — while schema shows none — is strong circumstantial evidence that lines up with independent survey data, but we’ll label it as evidence, not proof.

When are the results published?

We collect Wednesday and publish results Thursday with the full per-domain score sheet and charts, then a verdict Friday that judges each of the five predictions and turns the finding into an action list. This post is the locked design; nothing here changes once the data is in.

Sources

GeoParrot is a GEO Lab: we test what AI search actually cites and recommends, then publish the method and the misses. This is a Tuesday pre-registration in a four-part weekly arc — what AI SEO is, measured rather than asserted. Monday set the hypothesis; Wednesday we collect; Thursday the data; Friday the verdict.

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *

🦜 Follow GeoParrot: YouTubeX