Quick answer: This week we test the one claim Monday’s licensing fact-check left open: licensing didn’t buy general citations (paywalled partners got 0%, the open web 91.3%), but it might in fed verticals like local and shopping. So we’re pre-registering a test of three licensed/ingested partner domains — Yelp (330M reviews now inside ChatGPT), Reddit (licensed to Google), and Wikipedia (ubiquitously ingested) — against matched open-web comparators, across 12 queries in three buckets (general, local, shopping) on three engines. Our bet, locked below: Yelp spikes only in local and mostly on ChatGPT; Reddit is cited broadly; Wikipedia is a flat baseline; and the open web still holds the majority everywhere. If licensing were a general citation lever, partner domains would dominate all three buckets. We don’t think they will. Data collection starts Wednesday.
Experiment design — published September 8, 2026. This is a pre-registration post in the GEO Lab: we publish the hypothesis, the exact query set, the partner-vs-comparator scheme, and five locked predictions now, then collect and open the data later this week. Planting the goalposts in public first is the whole discipline.
What loose end is this experiment closing?
Monday we fact-checked the hot take that “a licensing deal is the new ranking factor.” The citation data said no: a July 2026 study found hard-paywalled and metered publishers — including the Financial Times, one of OpenAI’s own partners — earned 0% of AI citations, while the open, crawlable web captured 91.3%. A signed contract governs ingestion and legal cover; it doesn’t make your page readable when the engine assembles an answer.
But we were careful to flag the one scenario where licensing could still matter for visibility: vertical feed integrations. When an engine wires a licensed feed directly into an answer — Yelp’s reviews into ChatGPT’s local results, product catalogs into shopping answers — that partner’s data can crowd out the open web for those specific queries, because the engine prefers a clean, fresh, structured feed over crawling a dozen sites. That’s a claim that should be measured, not assumed. So this week we measure it in our own queries.
The stakes are practical. If licensing lifts citations generally, then “get on a licensed platform” becomes a real GEO tactic. If the lift is confined to fed verticals — and different per engine — then the advice collapses to something much narrower: get your structured data right if you’re a local or e-commerce business, and otherwise ignore the licensing headlines entirely. We expect the second answer, in line with our running finding that there is no universal GEO strategy.
What exactly are we measuring — and against what?
The metric is citation share: for each query on each engine, we pull the distinct domains the engine cites, then compute what fraction belong to a partner domain versus a matched open-web comparator of the same type. The comparison is what keeps this honest — we’re not asking “does Reddit get cited a lot” (it does), we’re asking “does a licensed partner out-cite a like-for-like source that has no deal.”
| Partner domain (licensed / ingested) | What it is | Matched open-web comparators |
|---|---|---|
| Yelp | Local directory, feed licensed to ChatGPT | Tripadvisor, other local directories/blogs |
| Community, data licensed to Google | Quora, Stack Exchange, niche forums | |
| Wikipedia | Reference, ubiquitously ingested | Britannica, official docs, editorial explainers |
We split the queries into three buckets, because the whole hypothesis is that licensing behaves differently by vertical:
- General (4): best CRM for small business, best project management software, best email marketing platform, best VPN. No licensed feed powers these — the control bucket.
- Local (4): best coffee shop in Austin, best dentist in Seattle, best plumber in Denver, best gym in Chicago. This is where Yelp’s feed lives.
- Shopping (4): best robot vacuum, best running shoes for beginners, best office chair, best air purifier. This is where product feeds live.
Twelve queries × three engines — ChatGPT (web search), Perplexity, and Google AI Mode — collected logged-out in a single snapshot, excluding each engine’s self-citations and Google’s own properties, the same exclusions we’ve used all series. Domains are classified into partner / matched-comparator / other-open-web, two-pass, and the full per-query table ships with Thursday’s results.
What five predictions are we locking before collection?
Pre-registered and numbered, each falsifiable. If the data contradicts one, Thursday’s post says so plainly.
- P1 — Yelp spikes only in local. Yelp’s citation share in the local bucket is at least 3× its share in the general bucket. A licensed feed lifts its vertical, not the whole board.
- P2 — Yelp’s lift is engine-specific. Yelp’s local citation share is at least 2× higher on ChatGPT than on Google AI Mode, because OpenAI holds the Yelp feed and Google does not. The licensing effect should not be uniform across engines.
- P3 — Reddit is the broad one. Reddit is cited in both general and vertical buckets and lands among the top-3 most-cited domains overall — community licensing generalizes where Yelp’s vertical feed does not.
- P4 — Wikipedia is a flat baseline. Wikipedia appears across buckets but its share does not spike in any single vertical; its cross-bucket variance is lower than Yelp’s. It’s ambient reference, not a fed vertical.
- P5 — The open web still wins the majority. Summed across all buckets, non-partner open-web domains hold ≥60% of total citation share. Licensing moves the vertical margins; it does not take over the answer.
Together these predictions describe a specific shape: licensing shows up as narrow, vertical, engine-specific feed presence — not a general citation lever. That’s the opposite of the “licensing = ranking factor” narrative. If P1–P5 hold, the headline take is wrong in exactly the way Monday argued. If Yelp turns up everywhere, or the open web falls below half, we were wrong and we’ll say it.
What could make this experiment misleading — and how do we stay honest?
Four limits, named in advance. One: it’s a single logged-out snapshot; AI citations drift week to week, so read this as a photograph, not a constant. Two: “citation share” is sensitive to how many domains an answer happens to cite — a bucket where answers cite fewer sources inflates each domain’s share, so we’ll report raw counts alongside shares. Three: ChatGPT rate-limits during collection, so its slice may be the thinnest and P2 will be read as directional if we can’t fully populate it. Four: local queries are US-city-specific and results vary by inferred location; we fix the cities in advance and note that the local picture may differ elsewhere.
We design in the open — same as the rest of the GEO Lab — because the licensing story is exactly the kind of headline that turns into confident, untested advice. “Get on the licensed platform” sounds actionable. This week we find out whether the engines’ actual citations back it or, as we suspect, quietly confine it to one or two verticals. Predictions are locked. Wednesday we collect; Thursday we open the data.
Frequently asked questions
What is this experiment actually testing?
Whether three licensed or heavily-ingested partner domains — Yelp, Reddit, and Wikipedia — get cited more than comparable open-web sources of the same type, and whether any advantage is general or confined to fed verticals like local and shopping. It’s the measured follow-up to Monday’s finding that licensing buys ingestion, not citations.
Why compare each partner against a “matched” open-web source?
Because Reddit, Yelp, and Wikipedia are cited partly for their format, not just their deals. Comparing Yelp to another local directory, Reddit to another forum, and Wikipedia to another reference site isolates the licensing/feed relationship from the page type — so if a partner out-cites its like-for-like comparator, the deal is a plausible cause rather than “forums just get cited a lot.”
Why pre-register the predictions instead of just reporting results?
Open-ended “here’s what we found” analysis lets you rationalize any pattern after the fact. Publishing five numbered predictions before collecting makes the test falsifiable: if Yelp shows up everywhere, or the open web drops below half, the record already says we were wrong. It’s the same discipline we use for every GEO Lab run.
Does this mean I should try to get my business “licensed” by an AI company?
No. Licensing deals are enterprise transactions offered to a short list of brand-name platforms with negotiating leverage — not something a typical site can obtain or needs. If you’re in a fed vertical (local, e-commerce), the practical lever is the eligibility layer: accurate, structured listings, reviews, and product data that a licensed feed can surface. Everyone else should stay crawlable and be the cleanest source in their category.
When are the results published?
We collect Wednesday and publish results Thursday with the full per-query table and charts, then a verdict Friday that judges each of the five predictions and turns the finding into a per-engine, per-vertical action list. This post is the locked design; nothing here changes once data is in.
Sources
- AI Is Buying Its Sources. Does a Licensing Deal Decide Who Gets Cited? — this week’s Monday fact-check (0% paywalled vs 91.3% open web; the Yelp–OpenAI and OpenAI publisher deals).
- Cross-engine citation results — the top ~15 domains account for roughly 68% of AI citations (Reddit, Wikipedia, YouTube lead).
- Reddit licenses to Google while suing scrapers — the community-licensing playbook this experiment tests.
- AI product discovery and shopping feeds and the share-of-voice vanity trap.
- GeoParrot GEO Lab — pre-registered design; collection to follow across ChatGPT (web search), Perplexity, and Google AI Mode, logged-out, single snapshot.

Leave a Reply