The verdict: ⚠️ Mostly rejected. Monday’s fact-check made two claims. The floor claim — “schema is table stakes, not a growth lever” — held: schema separated AI-cited domains from ignored ones by just 11.3 points, under our 15-point threshold. The lever claim — “first-party test data is the moat; publish it and AI cites you” — failed, and in the opposite direction. Across 148 domains scored blind on a pre-registered rubric, the sites AI cited averaged 0.45 on first-party test data while the sites it ignored averaged 0.69 (0–2 scale). The most-tested sites were the least-cited. So the honest ruling: the observation that AI leans on independent review outlets is real, but the causal lever Monday proposed — your original test data as an engineerable dial — did not survive its own experiment. Below: why it’s “conditional” and not a flat no, what to actually do per engine, and next week’s test.
Friday verdict — published September 18, 2026, closing this week’s GEO Lab arc. Monday we claimed schema is the floor and first-party test data is the moat. Tuesday we pre-registered four suspects and five falsifiable predictions. Wednesday we collected, Thursday we reported the numbers. Today we rule on the claim that started the week — including the parts that embarrassed us.
What’s the verdict on Monday’s claim?
Monday bundled four sub-claims into one confident thesis. Judged against Tuesday’s locked predictions, they split cleanly — two survived, two didn’t, and the one that didn’t is the one everybody would act on. Here’s the ruling, claim by claim:
| Monday’s sub-claim | Verdict | What the experiment showed |
|---|---|---|
| Schema is a floor, not a lever | ✅ Confirmed | Schema gap between cited and not-cited domains was 11.3pp — under our 15pp threshold (P2). Structure is close to a line everyone clears. |
| Brand size alone doesn’t buy a citation | ✅ Confirmed | 68 of 98 cited domains sat in the low/mid brand-demand tier; 19 high-demand sites (amazon, homedepot, salesforce, consumerreports) were cited zero times (P3). |
| AI defaults to independent test labs | ⚠️ Directionally true, mechanism wrong | The cited set really was RTINGS, Vacuum Wars, HouseFresh, SoundGuys, Security.org — but they did not win on the test-data score. The pattern is real; the proposed cause isn’t what separates them. |
| First-party test data is THE lever — publish it and you get cited | ❌ Rejected | Cited domains averaged 0.45 on test data vs 0.69 for the sites AI ignored — lower, not higher, and reversed from the pre-registered prediction (P1). This was the load-bearing claim. |
Net ruling: mostly rejected. The half of Monday’s advice that costs nothing to believe (schema is hygiene, brand size isn’t destiny) held up. The half you’d actually spend money on — “commission a test, publish the data, harvest the citations” — is the half our own pre-registered experiment knocked down. That’s an uncomfortable result to publish four days after asserting the opposite. We’re publishing it anyway, because the entire point of pre-registering predictions is that you don’t get to quietly delete the ones that miss.
Why “conditional” and not a flat “test data is useless”?
Because there’s one honest wrinkle we won’t hide behind, and it cuts against over-reading the reversal. Our comparator group — the “not cited” baseline — was built from Google’s organic top 10 for each query. That pool skews toward high-authority, well-established sites by construction. So if AI is citing a somewhat different, sometimes smaller-and-newer layer of the web than Google’s organic algorithm rewards, our comparators would look higher-scoring on almost everything — including test data — regardless of what actually drives citation. The fact that every one of the four suspects ran the same negative direction, with brand demand the widest gap, is exactly the fingerprint you’d expect from that bias.
So we hold two things at once, and the distinction is the whole verdict:
- Well supported: “First-party test data, on its own, is not the dial that separates AI-cited domains from the sites that rank but don’t get cited.” A prediction that specific, failing that directionally, is a real finding whatever the comparator method.
- Not supported (and we won’t claim it): “Test data or authority is irrelevant to AI citation.” That would need a hand-matched comparator group — rival test site vs. manufacturer page vs. general blog — which we didn’t build this round. We flagged the shortcut in Thursday’s results and we’re not laundering it into a bigger claim now.
That’s why the verdict is conditional. Monday’s picture — AI reaching past manufacturer pages for independent reviewers who tested the thing — is what we actually saw in the citation logs. Monday’s prescription — “so publish your own test data and you’ll get cited” — is the part that failed. You can be right about the room and wrong about the door.
What should you actually do, per engine, given this?
The result kills a shortcut, not a strategy. Here’s the checklist we’d hand a team on Monday morning, split the way the data splits — because ChatGPT and Google AI Mode barely cite the same domains, and shopping, software, and local behaved differently.
ChatGPT (web search)
- Get tested by a source AI trusts, don’t just test yourself. The cited set was independent outlets (RTINGS, SoundGuys, PCMag), not brand pages. This is the earned-beats-owned pattern again: a third party’s coverage of your product outperformed a brand’s own “best of” page by roughly 6× in citation share. Earn the review; don’t only publish the spec sheet.
- Clear the floor once, then stop. Schema, parseable specs, an FAQ — do them because they make you legible, not because they’ll win the citation. The 11.3pp schema gap confirms they’re table stakes, not a lever.
Google AI Mode
- Stop assuming your organic rank transfers. Nineteen sites that ranked in Google’s own organic top 10 — including amazon.com and consumerreports.org — were cited zero times by AI Mode for the same query. Ranking #1 is not a citation. Track them as separate scoreboards.
- For local intent, specificity beats size. A long tail of small individual business pages surfaced in local queries where big directories didn’t. Hyper-relevant, single-location pages punched above their authority — the opposite of chasing domain rating.
Both engines, category by category
- Shopping “best X”: independent labs plus YouTube own it. If you sell the product, your page is unlikely to be the cited source — budget for earned coverage, not a self-published scorecard.
- Measure per engine and per category, never in aggregate. The single biggest through-line of this whole series stands: there is no universal GEO lever. A blended “AI visibility score” would have hidden every finding above.
- Stop treating the demo as causal. “We added schema and a test, then appeared in AI” is a before/after screenshot, not a controlled result. In our data the highest test-data scores belonged to the sites that were not cited.
What are we testing next week?
This verdict leaves one clean question on the table, and it’s the one that would actually resolve the reversal. If the most-tested sites aren’t the most-cited, maybe the signal isn’t test data at all — it’s independence. Next week we build the hand-matched comparator we skipped this round: for the same product, we’ll pit the manufacturer’s own test page (they measured it, they have first-party data) against the independent reviewer who tested the same product — matched pair, same query, scored blind. If AI reaches for the independent reviewer even when the manufacturer holds equal or better first-party data, then the lever was never “who has test data” — it’s “who has no incentive to lie.” That also finally lets us score the cited-vs-recommended gap (P5) we couldn’t test this run. Monday opens it; follow the GEO Lab to see whether independence survives a fair fight.
Frequently asked questions
So does publishing first-party test data get you cited by AI or not?
Not as a standalone lever, based on our pre-registered experiment. We predicted AI-cited domains would score at least 2× higher on first-party test data than matched comparators; instead they scored lower — 0.45 vs 0.69 on a 0–2 scale. Test data is a genuine asset and the cited sites were mostly independent testers, but “publish a test, harvest the citation” is a causal claim our own data contradicted. Getting covered by a source AI already trusts beat self-publishing.
What’s the final verdict on Monday’s fact-check?
Mostly rejected. Two sub-claims held (schema is a floor, not a lever; brand size alone doesn’t buy a citation). The load-bearing sub-claim — first-party test data is the engineerable lever — failed in the opposite direction and is rejected. The observation that AI defaults to independent test labs is directionally true, but the mechanism Monday proposed for it did not survive.
If test data isn’t the lever, why did you say cited sites were RTINGS and SoundGuys?
Both are true. The sites AI cited were overwhelmingly independent review and testing outlets — that’s what the citation logs show. But when we scored those same sites on a locked first-party-test-data rubric, they did not out-score the ranking sites AI ignored. So “AI cites independent testers” is an accurate description; “and it’s their test-data score that causes it” is the part the experiment rejected. The likely real signal — independence, not the data itself — is what we test next week.
Should I stop adding schema to my pages?
No — add it once as eligibility hygiene, then stop optimizing it as if it were a growth lever. The schema gap between cited and not-cited domains was only 11.3 points, confirming it’s close to a floor everyone clears. Structured data makes your page legible to a model; it doesn’t decide whether the model picks you as the source.
Why publish a verdict that contradicts your own Monday post?
Because we locked five falsifiable predictions on Tuesday before scoring anything, precisely so we couldn’t rewrite them after the numbers came in. When the load-bearing prediction reversed, the discipline is to report the reversal, not reframe the claim. A GEO Lab that only publishes its wins is a marketing blog, not a lab.
Sources
- Schema Won’t Get You Cited. First-Party Test Data Will — Monday’s fact-check, the claim under judgment.
- Four Suspects, Locked Before We Look — Tuesday’s pre-registered rubric and five predictions.
- Cited Sites Scored Lower On All Four Suspects — Thursday’s full 148-domain results and methodology limits.
- Earned beats owned in AI citations — the third-party-coverage pattern behind this week’s checklist.
- There is no universal GEO strategy — the running finding this verdict extends.
- GeoParrot GEO Lab — exp14: 15 buyer-intent queries × ChatGPT (web search) + Google AI Mode, logged-out single snapshot; 98 cited + 56 comparator domains (148 unique), scored 0–2 on four pre-registered suspects.
GeoParrot is a GEO Lab: we test what AI search actually cites and recommends, then publish the method and the misses — including the weeks our own Monday claim doesn’t survive to Friday. This closes a four-part weekly arc on what AI SEO is, measured rather than asserted.

Leave a Reply