Quick answer: We scored 24 buyer-intent answers (12 “best [tool] 2026” questions × Gemini and Perplexity) to test the popular advice: publish your own best-of listicle, rank yourself #1, and win the AI recommendation. It fails at both ends. A major category vendor’s own listicle was cited in only 8.3% of answers (2 of 24), the #1 recommended brand’s own site was the cited source in just 4.2% (1 of 24), and in the two cases a vendor did cite its own list, it lost the top pick to a rival both times (Salesforce → Pipedrive, Wrike → ClickUp). Meanwhile the two engines agreed on the exact #1 pick in 6 of 12 categories and shared overlapping top picks in all 12 — recommendations are anchored to category reputation, not to who ranked themselves first.
On Monday we flagged the self-ranking trap — the advice to publish your own “Top 10” and put yourself at #1, set against Lily Ray’s finding that 69% of AI Overview citations of self-promotional listicles recommend a competitor in the list. On Tuesday we pre-registered a method with a novel two-axis rubric — cited (your page is the linked source) vs recommended (the answer’s prose actually picks you) — and locked four predictions before collecting a single answer. This post is the raw result, graded against what we said would happen.
Does publishing your own listicle even get you cited?
Before you can be “cited but not recommended,” your page has to be cited at all. So the first question is upstream of the whole self-ranking thesis: when someone asks an engine “best CRM software 2026,” does it reach for HubSpot’s or Salesforce’s own “best CRM” roundup? Overwhelmingly, no.

Both engines cite a mix — independent review media (PCMag, TechRadar, G2, ZDNet, Venture Harbour, Security.org) sit alongside plenty of vendor-owned domains. But the specific play the advice recommends — a major category incumbent citing its own self-promotional best-of listicle — barely happens:
- Self-serving citation rate: 8.3% — a pre-registered category vendor’s own listicle domain was cited in just 2 of 24 answers (both on Perplexity: salesforce.com, wrike.com). Gemini cited a big incumbent’s own list zero times in 12 answers.
- The vendor-owned domains engines do cite are mostly niche tools’ blog content (Botpress, Gleap, Netcore, Helpmonks) — cited for a roundup, not because they’re the household name in the category.
So the first leg of the advice is already shaky. You cannot reliably manufacture a citation by publishing your own “best of” — the engines lean on third-party reviewers and a long tail of other people’s content.
When your page is cited, does the answer recommend you?
This is the measurement the method was built for, and it’s the sharpest number in the run. We checked, for every answer, whether the brand the prose actually recommended #1 traced back to that brand’s own site being the cited source.

The #1 recommended brand’s own site was the cited source in only 4.2% of answers (1 of 24). In other words, when an engine says “go with Zendesk” or “Mullvad is the privacy leader,” that recommendation almost never rests on Zendesk’s or Mullvad’s own page — it rests on an independent reviewer. Getting cited and getting recommended are decoupled:
- Perplexity, CRM (Q1): cited
salesforce.comas a source — then recommended Pipedrive, Zoho, and HubSpot as the top picks, with Salesforce demoted to “also a strong choice.” - Perplexity, project management (Q3): cited
wrike.comas a source — then made ClickUp the #1 pick and listed Wrike last.
Two out of two self-citing vendors were cited and still lost the recommendation to a rival. That’s exactly the “citation ≠ recommendation” gap Lily Ray measured on Google’s AI Overviews — reproduced here, transparently, on Gemini and Perplexity.
If not self-ranking, what decides the recommendation?
Category reputation. When we compared the two engines’ picks head-to-head across the 12 categories, the recommendations were strikingly stable — which is the opposite of what you’d see if a well-placed self-ranked listicle could swing the answer.

- Same #1 pick in 6 of 12 categories — help desk (Zendesk), live chat (LiveChat), marketing automation (HubSpot), survey (SurveyMonkey), accounting (QuickBooks), and VPN (Mullvad). Two independent engines, identical top pick.
- Overlapping top picks in all 12 of 12 — even where the #1 differed (CRM, email, PM, passwords, website builders, AI writing), the two engines’ shortlists shared at least one recommended brand.
Recommendations track who the category’s independent reviewers already crown, and both engines converge on the same well-known names. That consensus is not something you edit by publishing a page that ranks yourself first.
How the pre-registered predictions scored
We locked four predictions on Tuesday. Here’s the honest grade — with the caveat that only 2 of 24 answers produced a self-serving citation, so anything about that subset is directional, not statistical:
| Prediction (from the method post) | Result | Grade |
|---|---|---|
| #1. When a vendor cites its own listicle, the publisher itself is recommended <35% of the time | 0 of 2 self-citing vendors won the top pick (0%) | ✅ directional (n=2) |
| #2. Third-party listicle citations convert to a recommendation more than self-published ones | The #1 rec’s own site was the source in just 1/24; recommendations traced to reviewers, not self-published pages | ⚠️ supported, under-powered |
| #3. The recommendation matches the publisher’s own #1 <50% of the time | 0 of 2 self-citing vendors’ self-#1 matched the engine’s pick | ✅ directional (n=2) |
| #4. No engine recommends the self-publishing vendor a majority of the time | Confirmed — self-serving citations were 8.3% and none became the recommendation | ✅ |
The predictions held in direction, but the run also surfaced something we didn’t predict and that matters more: the self-ranking play mostly fails one step earlier, at the citation itself (8.3%), and recommendations turned out to be far more reputation-locked across engines than a per-engine tuning story would suggest.
What we can’t claim yet (the honest limits)
- Two engines, not four. ChatGPT sat behind a logged-out login wall the whole run and Google’s AI Mode (
udm=50) aborted every request from our environment. This is Gemini + Perplexity only. Predictions #3 and #4 wanted per-engine breadth we couldn’t get. - The self-serving subset is n=2. Only two answers cited a big vendor’s own listicle, so “self-citers lose the recommendation” is a clean, consistent anecdote — not a rate you should quote to three decimals.
- Perplexity throttled its source panel. Nine of its 12 answers exposed sources; the other three returned prose without a citation list, so the provenance counts for Perplexity are drawn from those nine.
- Scoring of “recommended” was a two-pass read of the answer prose against a rubric frozen before collection. The pattern is strong; treat the exact percentages as directional given the sample.
What this means if you’re doing GEO
The takeaway is blunt: you can’t self-rank your way into an AI recommendation, and you probably can’t even self-rank your way into a citation. What actually moves the needle for buyer-intent queries:
- Earn placement in the third-party lists engines already trust — PCMag, TechRadar, G2, Capterra, category review media. That’s where citations and recommendations come from. (See our breakdown of what page format engines actually cite: 66% of citations are ranked best-of listicles.)
- Build the category reputation the reviewers reward. Recommendations converge on the same names across engines because those names are consensus picks — the durable GEO asset is being a genuine category leader third parties keep listing, not a self-published #1.
- If you publish your own roundup, be honestly comparative. A credible list that ranks rivals fairly can still get cited for its content; a transparent self-serving “we’re #1” page is the thing engines route around — and, per Lily Ray, the thing Google now suppresses in organic too.
- Don’t over-tune per engine yet — for these categories the two engines’ recommendations were remarkably aligned. Win the shared, reputation-driven layer first (echoing our cross-engine overlap experiment).
On Friday we’ll close the week with the verdict — does self-ranking ever pay, and a per-engine self-ranking playbook — plus the throttle and small-sample caveats folded into the final call.
Frequently asked questions
Does publishing my own “best [category]” listicle get me cited by AI?
Rarely, in our run. A major category vendor’s own best-of listicle was the cited source in only 8.3% of 24 buyer-intent answers (Gemini: 0 of 12). Engines leaned on independent review media and a long tail of other sites instead.
Is being cited the same as being recommended?
No — that’s the core finding. The #1 recommended brand’s own site was the cited source in just 4.2% of answers. In both cases where a vendor cited its own listicle, the engine recommended a competitor (Salesforce → Pipedrive, Wrike → ClickUp).
Why only Gemini and Perplexity?
ChatGPT required a logged-in session our unattended collector couldn’t complete, and Google’s AI Mode endpoint aborted every request from our environment. Rather than fabricate coverage, we report the two engines we could collect cleanly and flag the gap.
If self-ranking doesn’t work, what does?
Getting into the third-party lists engines cite (PCMag, TechRadar, G2, Capterra) and building genuine category reputation. Recommendations were reputation-anchored — the two engines agreed on the exact #1 pick in half the categories and shared top picks in all of them.
Can I reproduce this?
Yes. The 12 buyer-intent questions, the per-engine raw captures (URLs, cited domains, and answer prose), the frozen scoring rubric, and the aggregation script all live in our GEO Lab exp5 folder. The predictions were pre-registered in Tuesday’s method post before any answer was collected.

Leave a Reply