Does a Reddit Mention Actually Cause an AI Citation? Our Pre-Registered Method

Quick answer: To test whether a Reddit mention causes AI citations (rather than just correlating with brands that were already citable), we’re running three convergent tests: (1) provenance — when an engine cites a brand, does it actually link a reddit.com URL, or is Reddit just coincidentally present? (2) matched pairs — compare same-category brands with similar domain authority but different Reddit footprints, so authority can’t explain the gap; and (3) a slow, organic intervention pilot — make one genuine, disclosed Reddit contribution for a low-citation brand and re-measure weeks later. We’ve pre-registered four predictions before collecting a single data point. No fake upvotes, no manufactured threads — this post is the method, published before the results so you can hold us to it.

Yesterday we showed that the “Reddit is AI’s #1 source” stat falls apart the moment you line up the studies — Semrush says 40%, Bluefish says YouTube just overtook it, and Perplexity’s Reddit share is reported anywhere from 6.7% to 46.7%. But underneath that noise sits a bigger question that no published study has cleanly answered: when a brand gets mentioned on Reddit and then gets cited by AI, did the Reddit thread cause the citation — or do AI engines just happen to cite brands that were already strong enough to be discussed on Reddit in the first place?

That gap between correlation and causation is the whole game for anyone deciding to spend time on Reddit. So this week at GEO Lab we’re testing it ourselves — fully in the open, with the same pre-registration discipline we used for our 4-engine overlap experiment. Here’s exactly how, before we touch the data.

What question are we actually answering?

One testable claim, stated plainly:

Hypothesis (H1): A brand’s presence in a relevant Reddit thread increases the probability that AI engines cite that brand for buyer-intent queries — a real causal lift, beyond what the brand’s general authority would predict.
Null hypothesis (H0): Reddit presence only correlates. Engines cite brands that were already citable (strong site, backlinks, reviews); the Reddit thread is a symptom of being citable, not a cause of the citation.

The trap here is that the easy evidence — “this cited brand also has a Reddit thread” — is exactly what you’d see under both hypotheses. Correlation looks identical to cause until you control for the thing doing the confounding: authority. That’s what our design is built to separate.

Why is causation so hard to prove here?

You can’t run a clean randomized trial on Reddit. The honest constraints, stated up front:

  • Confounding by authority. Brands big enough to be debated on Reddit usually also have the backlinks, reviews, and content that independently earn AI citations. Reddit and citation-worthiness rise together.
  • You can’t ethically manufacture the cause. Astroturfing — fake accounts, paid upvotes, planted threads — would both violate Reddit’s rules and invalidate the experiment, because a manufactured thread isn’t the organic signal engines are trained to trust. So we won’t do it. Zero fake upvotes.
  • Indexing lag. Even a genuine Reddit contribution takes days-to-weeks to be crawled, indexed, and reflected in an engine’s retrieval. A same-day re-query proves nothing.

No single test beats all three problems. So instead of pretending one clean experiment exists, we triangulate: three tests that fail in different directions, and only agree if there’s a real effect.

What are the three tests?

Test What it isolates What a “Reddit causes it” result looks like
A · Provenance Is Reddit the cited source, or just present? Engines frequently link the actual reddit.com URL — not just a brand that also has a thread
B · Matched pairs Effect of Reddit after removing authority The higher-Reddit brand gets cited more than its similar-authority twin
C · Intervention pilot Actual before/after of adding Reddit presence Citation probability rises after a genuine Reddit contribution is indexed

Test A — Provenance: cited source vs. coincidental presence

For every brand an engine cites across our buyer-intent question set, we record how it was supported. Three buckets:

  • Reddit-as-source: the engine’s citation links an actual reddit.com thread.
  • Reddit-adjacent: the brand is cited via its own domain (or a review site), and a relevant Reddit thread exists but the engine didn’t link it.
  • No-Reddit: the brand is cited and has essentially no relevant Reddit footprint.

If “get on Reddit” causes citations, the first bucket should be large. If most cited brands are Reddit-adjacent or No-Reddit, then Reddit’s role is, at best, indirect — which already deflates a lot of advice.

Test B — Matched pairs: controlling for authority

This is the core of the causal claim. We pick pairs of competing brands in the same category with roughly equal domain authority and backlink profiles, but very different Reddit footprints (one heavily discussed, one barely). Same category, same buyer questions, similar authority — the main thing that differs is Reddit. If the high-Reddit brand is cited materially more often, authority can’t explain it, and Reddit becomes the leading suspect. If the gap vanishes once authority is matched, that’s strong evidence the “Reddit effect” was really an authority effect all along — exactly the kind of divergence we found when four engines cited almost entirely different sources.

Test C — The intervention pilot (the real causal probe, run slowly)

The gold standard is to change the cause and watch the effect. So we’ll identify a small number of genuinely low-citation brands/topics, make one honest, on-topic, disclosed Reddit contribution that a real user would find useful, then re-query the engines after a multi-week indexing window. Pre-citation vs. post-citation = a directional causal estimate. We’re flagging this clearly: it’s a tiny sample, it’s slow, and it can only run on its own two-week-plus timeline — so its result lands in a later GEO Lab post, not this week’s verdict. Tests A and B carry this week.

How do we collect the data without fooling ourselves?

Risk How we control for it
Authority confound Test B matches brands on domain authority / backlinks before comparing
“Present ≠ cited” illusion Test A separates an actual reddit.com link from a thread that merely exists
Personalized results Logged-out / fresh sessions, no memory, neutral account, same region
Run-to-run variance Each question asked multiple times per engine; we report the spread, not one lucky pull
Engines blocking us We pace for throttling — Perplexity and Gemini systematically throttled our last run, so coverage gaps get reported, never hidden

Collection uses a browser agent driving each engine’s real interface across ChatGPT Search, Perplexity, Google AI Mode and Gemini — the same setup as our first 3-engine test, where we learned which engines throttle and which hide their sources in side panels.

Our predictions — locked in before the data

Pre-registration is what separates an experiment from a story told afterward. On the record, before collecting a single citation:

  1. Under 25% of brand citations will be Reddit-as-source (an actual reddit.com link). Most “Reddit effect” will be indirect, not a direct citation.
  2. In matched pairs, the high-Reddit brand will be cited more often — but the gap will be smaller than the raw correlation implies, because authority explains a big chunk.
  3. Engines will split: Perplexity and Google AI Mode will link Reddit directly more than ChatGPT and Gemini, which tend to name brands without linking the thread.
  4. The honest verdict will land on Reddit as a contributing factor, not a sole cause — it raises the odds of citation but doesn’t manufacture one on its own.

If the data contradicts any of these, we’ll say so on Friday. That’s the entire point of writing them down where you can see them.

Next at GEO Lab: we collect provenance and matched-pair data midweek and publish the raw results later this week — every number against these four predictions — then deliver the verdict Friday: does a Reddit mention actually cause AI citations, or has the whole industry been selling correlation as cause? Start with the claim →

FAQ

Does getting mentioned on Reddit cause AI to cite you?
That’s exactly what no published study has cleanly proven — most show correlation. A brand on Reddit is usually also a brand with strong authority, and either could drive the citation. Our experiment is built to separate the two by matching brands on authority and checking whether Reddit is the actual cited source.

Why can’t you just run a normal A/B test on Reddit?
Because you can’t ethically manufacture the “cause.” Planting threads or buying upvotes violates Reddit’s rules and would invalidate the result, since engines are tuned to trust organic signals. So we triangulate with three observational and one slow, organic intervention test instead.

Why pre-register the predictions?
Because it’s easy to look at results and craft a tidy explanation afterward. Writing predictions first makes the test falsifiable — if the data contradicts us, you’ll know we didn’t move the goalposts.

When do we get the answer?
The provenance and matched-pair results land later this week, with a verdict Friday. The slower intervention pilot reports in a later post, because genuine Reddit indexing takes weeks.

Sources

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *