Does Ranking Yourself #1 in Your Own Listicle Win the AI Recommendation? Our Pre-Registered Method

Quick answer: Monday we confirmed the uncomfortable half of the “publish your own Top 10” advice: self-serving listicles do get cited (vendors earn 59.5% of buyer-intent citations, 62% of them from their own best-of pages) — but a 100-query study found the AI Overview handed the recommendation to a listed competitor 69% of the time. This week we test that gap ourselves, across four engines. Our pre-registered method has three parts: (1) a self-serving-listicle audit — when an engine cites a vendor’s own best-of page, does it recommend the publisher or a competitor from that same list?; (2) a self-published-vs-third-party head-to-head — do independent listicles (TechRadar, PCMag, directories) convert a citation into a recommendation more reliably than a vendor’s own list?; and (3) a self-ranking-position test — when a vendor crowns itself #1, does the engine actually pick #1, or re-rank on its own signals? We’ve locked four predictions before collecting a single answer. This post is the method, published before the results, so you can hold us to it.

Yesterday’s fact-check opened this week’s question: self-published listicles get cited, but a citation is not a recommendation — and for self-serving lists specifically, Lily Ray’s data suggests the recommendation often goes to a rival you listed. That’s a claim we can test directly, so this week we do. Today, before we collect a single answer, here’s exactly how — with the same pre-registration discipline we used for our Reddit causation method and our cross-engine overlap method.

What question are we actually answering?

One testable claim, stated plainly:

Hypothesis (H1): Ranking yourself #1 in your own best-of listicle earns a citation but not a recommendation. When AI engines cite a self-serving vendor listicle, they recommend the publishing brand a minority of the time and hand the recommendation to a listed competitor more often than not. Independent third-party listicles convert their citations into recommendations more reliably than self-published ones.
Null hypothesis (H0): A citation is a recommendation. When a vendor’s own listicle gets cited, the engine recommends that vendor about as often as any listed brand, self-assigned #1 position carries through, and the source of the list (own vs third-party) makes no difference to who gets recommended.

If H1 holds, “publish a Top 10 and put yourself first” is optimizing the cheap half of the funnel (the citation) while feeding the expensive half (the recommendation) to rivals — and the GEO play shifts to earning spots on lists engines trust. If H0 holds, self-ranking works exactly as the advice promises. The method below is built to tell those two worlds apart.

What’s the difference between “cited” and “recommended” — and how do we measure it?

This is the whole experiment, so we define it before we collect anything. Every AI answer to a buyer-intent query gets scored on two independent axes for each brand that appears:

Axis Definition (frozen before collection) How we detect it
Cited A page owned by, or ranking, the brand appears as a linked source in the answer The source URL is present in the citation list / footnotes / inline links
Recommended The answer names the brand as a pick — “best overall,” a lead recommendation, or a top entry the answer endorses in its own prose The brand appears in the answer’s synthesized recommendation text, not merely in a scraped list

The key move is that a brand can be cited but not recommended (its self-serving page is a source, yet the engine tells the user to pick someone else), recommended but not from its own page (a third-party list carries it), or both. For every cited self-serving listicle, we record which of the listed brands the answer actually recommends: the publisher, a listed competitor, or neither. That single classification is what turns Monday’s citation-share numbers into a recommendation-conversion rate.

What corpus are we collecting — and why a fresh pull this time?

Last week we deliberately reused our provenance corpus, because the question was about citation shares and the old data preserved every source URL. This week is different: measuring recommendation requires the answer’s synthesized prose — which brand it endorses — not just the list of source links. Our earlier corpora stored URLs and domains, not the recommendation verdict, so they can’t answer this. So we collect fresh (call it exp5): ~12 buyer-intent “best [category]” queries × 4 AI engines (ChatGPT with web search on, Perplexity, Google AI Mode, Gemini), each run three times, logged out, within a 48-hour window, capturing the full answer text and the cited sources. Categories are chosen where self-serving vendor listicles are known to exist (CRM, email marketing, project management, password managers, and similar), so both self-published and independent lists have a real chance to be cited and compared. The tradeoff, stated up front: it’s a small, hand-run sample — treat every result as “the pattern is strong,” not “these percentages to the decimal.” Where an engine throttles us — Perplexity gave usable sources for only 1 of 12 questions in our last provenance run — that engine is under-represented and we’ll say so, never average the gap away.

What are the three parts of the test?

Part What it measures What an “H1 is right” result looks like
1 · Self-serving-listicle audit When a vendor’s own best-of page is cited, who gets recommended The publisher is recommended a minority of the time; a listed competitor wins more often
2 · Self-published vs third-party Whether independent lists convert citations to recommendations better than self-published ones Third-party listicle citations lead to a recommendation of a listed brand at a higher rate
3 · Self-ranking-position test Whether being #1 in your own list wins the pick The recommended brand rarely matches the publisher’s self-assigned #1 — the engine re-ranks

Part 1 — The self-serving-listicle audit

For every answer, we first flag whether any cited source is a self-serving listicle — a vendor-owned “best [category]” page that ranks the publisher’s own product among the options (usually at or near #1). The rubric for “self-serving,” frozen before collection: the page is on the vendor’s own domain and lists that vendor as one of the ranked entries. For each such cited page, we record who the answer recommends — publisher, listed competitor, or neither — and compute the publisher-recommendation rate: of all cited self-serving listicles, what share resulted in the engine recommending the brand that published them. This is our multi-engine replication of Ray’s Google-only 69%-competitor finding, scored transparently rather than asserted.

Part 2 — Self-published vs third-party head-to-head

Type of list, not just type of page. For the same queries, we separate cited listicles into self-published (vendor-owned, ranking themselves) and independent third-party (TechRadar, PCMag, category directories — the domains engines actually reach for). For each group we measure the citation-to-recommendation conversion rate: when a list of that kind is cited, how often does the engine go on to recommend a brand from it in its own prose? The hypothesis is that engines treat an even-handed independent list as a trustworthy basis for a recommendation, while a transparent self-serving list gets scraped for coverage but not trusted for the pick. If both convert equally, the “own vs earned” distinction collapses and self-publishing is as good as placement.

Part 3 — The self-ranking-position test

This is the part that tests the advice head-on. When a self-serving listicle is cited and the answer recommends one of its listed brands, does the recommended brand match the publisher’s self-assigned #1? We record the publisher’s own ranking of the entries and compare it to the engine’s pick. If crowning yourself #1 works, the engine’s recommendation should usually be that #1 entry (the publisher). If the engine re-ranks on its own signals — reviews, mentions, comparison evidence — the recommended brand will frequently be a lower-ranked entry or a competitor, proving self-assigned order is decorative. This is the difference between “the list influences the answer” and “your position in the list influences the answer.”

How do we score without fooling ourselves?

Risk How we control for it
“Recommended” is subjective The cited/recommended rubric is written and frozen before collection; a brand counts as recommended only if it appears in the answer’s own endorsement prose, not merely in a scraped list
Motivated reading of ambiguous answers Two passes per answer; disagreements re-checked against the rubric, edge cases logged with the reason rather than silently forced
Run-to-run variance Three runs per query in a 48-hour window, logged out, identical wording; we report the distribution, not a single lucky pull
Cherry-picked categories Query list fixed before collection and published with the results; categories span several verticals, not one
Engine imbalance Per-engine coverage reported (Perplexity throttled to ~1/12 last run); we never average a gap away

The query list, every captured answer’s cited/recommended coding, and the per-listicle publisher-vs-competitor calls will ship as a downloadable table with the results — the same reproducibility standard as our cross-engine data, so anyone can re-run the counts and disagree with a specific call.

Our predictions — locked in before we collect a single answer

Pre-registration is what separates an experiment from a story told afterward. On the record, before collection:

  1. Self-serving listicles are cited but rarely recommend the publisher. Of cited self-serving listicles, the publishing brand is the recommended pick in under 35% of cases — a listed competitor or “neither” wins the majority, echoing Ray’s 69%-competitor result across more than one engine.
  2. Third-party lists convert better. Independent third-party listicles convert a citation into a recommendation of a listed brand at a higher rate than self-published listicles do. Engines trust earned lists as a basis for the pick more than self-crowned ones.
  3. Self-assigned #1 doesn’t carry. When a cited self-serving listicle leads to a recommendation, the recommended brand matches the publisher’s own #1 in under 50% of cases — the engine re-ranks on its own signals, so your position in your own list is decorative.
  4. No engine rescues self-ranking. On no single engine does the self-serving publisher get recommended a majority of the time. Link-transparent engines (Perplexity, Google AI Mode) surface the underlying competitor more visibly, but even brand-synthesizing engines (ChatGPT, Gemini) don’t reliably crown the publisher.

If the data contradicts any of these, we’ll say so on Friday. That’s the entire point of writing them down where you can see them.

Next at GEO Lab: we collect the four-engine answers midweek and publish the results Thursday — every number against these four predictions, with charts — then deliver Friday’s verdict and a self-ranking playbook: given what engines actually do with your own listicle, should you publish it, earn a spot on someone else’s, or both? Start with Monday’s fact-check →

FAQ

What are you actually testing this week?
Not whether self-published listicles get cited — Monday answered that (they do; vendors earn 59.5% of buyer-intent citations). This week tests whether that citation becomes a recommendation: when an engine cites your own best-of page, does it recommend you, or a competitor you listed?

How do you tell a citation from a recommendation?
A citation means your page is used as a source (its URL appears in the answer). A recommendation means the answer names your brand as a pick in its own endorsement prose. We score both axes independently for every brand in every answer, with a rubric frozen before collection, so a brand can be cited-but-not-recommended — which is exactly the trap we’re measuring.

Why collect fresh data instead of reusing last week’s corpus?
Because “recommended” lives in the answer’s synthesized text, and our earlier corpora only stored source URLs, not which brand the engine endorsed. Measuring recommendation requires capturing the full answer, so we collect a fresh ~12-query, 4-engine pull and publish the query list and codings with the results.

Isn’t judging “recommended” subjective?
Somewhat — which is why the rubric is frozen before collection, each answer is scored twice, ambiguous cases are logged with a reason, and the full coded dataset ships with the results so you can disagree with any specific call. Three runs per query also guard against a single lucky answer.

When do we get the answer?
The results and charts land Thursday, with the verdict and self-ranking playbook Friday.

Sources

Comments

Leave a Reply

Your email address will not be published. Required fields are marked *