TL;DR
- Reddit accounts for roughly 46% of Perplexity citations and ChatGPT leans heavily on Wikipedia; product and marketing pages together make up around 3% of cited URLs.
- This is not a ranking accident. It reflects training data overlap, licensing deals, structural answerability, and the way retrieval systems weight neutral, multi-perspective content.
- LLMs cite sources that resolve the question without the user having to filter sales language. Homepages do the opposite.
- You cannot out-Wikipedia Wikipedia. You can, however, build pages that look like reference material: declarative answers, primary data, named authors, and clear scope.
- The fastest wins: convert money pages into evidence pages, get cited on Reddit and Wikipedia (not just by them), and publish original data competitors will quote.
Most brand teams discover the same uncomfortable pattern the first time they run citation tracking: ChatGPT quotes Wikipedia, Perplexity quotes Reddit threads from 2019, and the carefully optimized homepage barely registers. This isn't a bug in the engines — it's a direct consequence of how retrieval-augmented generation chooses sources. Below is what the data shows, why it happens, and the realistic playbook for brand pages that want a seat at the table.
What the citation data actually shows
Across the largest public studies of AI citations, two domains dominate. Semrush's three-month analysis of ChatGPT, Google AI Mode, and Perplexity found Reddit and Wikipedia at the top across all three engines, with Reddit especially concentrated inside Perplexity. Profound's platform breakdown put Reddit at roughly 46.7% of Perplexity citations — nearly half of every source the engine surfaces.
ChatGPT's profile skews differently. Ahrefs' study of 9.6 million ChatGPT queries showed Wikipedia as the single most-cited domain, with reference sites, news publishers, and government domains rounding out the top of the list. Azoma's longitudinal tracking shows Wikipedia's share inside ChatGPT has grown, not shrunk, as the product matures.
The flip side: Neil Patel's analysis of 10,000 ChatGPT sessions found that product and marketing pages make up only a low single-digit percentage of cited URLs. The pages your team spends the most time on are the pages LLMs trust least.
Why LLMs prefer Reddit and Wikipedia
Four structural reasons, in rough order of weight:
1. Training data gravity. Both Reddit and Wikipedia were heavily represented in the pretraining corpora of every major frontier model. The model has stronger internal representations of these domains, and retrieval systems are tuned to confirm — not contradict — what the base model already "knows." OpenAI's licensing agreement with Reddit reinforces this for ChatGPT specifically.
2. Answer-shaped content. Wikipedia paragraphs are declarative, scoped, and start with the definition. Reddit threads contain the exact question a user typed, followed by ranked human answers. Both formats map cleanly onto the chunk-and-rank step inside RAG pipelines. A homepage hero that says "Reimagine your workflow" maps onto nothing.
3. Neutrality signals. Retrieval systems penalize promotional language at the embedding and reranking layers. Wikipedia enforces a neutral point of view by policy. Reddit answers, even biased ones, are framed as opinions among other opinions — which engines can present as "users report." A vendor claiming the same thing about itself reads as a conflict of interest.
4. Freshness and breadth on long-tail queries. For any question that hasn't been formally documented, Reddit is often the only place a real answer exists. Perplexity's heavy Reddit weighting reflects this: it's optimizing for queries where curated sources simply don't have coverage.
Why your homepage loses
Homepages and product pages fail AI retrieval on almost every axis. They're written for a buyer mid-funnel, not for someone asking a factual question. They bury specifics under value propositions. They rarely cite primary sources, rarely name an author, and rarely answer a question in the first sentence. The semantic embedding of a typical SaaS homepage is closer to other SaaS homepages than it is to any user query.
There's also a structural issue: homepages tend to be the most-linked but least-informative page on a domain. Backlinks help discovery, but retrieval is decided at the chunk level. A 60-word hero section, however authoritative the domain, will lose to a 300-word Wikipedia paragraph every time.
What to do instead
You will not out-rank Wikipedia on its own ground. The realistic plays:
Convert money pages into evidence pages. Add a "How it works" section with concrete mechanics, a named author with credentials, dated last-updated stamps, and inline citations to primary sources. The page can still convert — but it has to earn the citation first by looking like reference material.
Publish original data. The fastest way onto Reddit, Wikipedia, and downstream AI citations is to be the primary source someone else quotes. Surveys, benchmarks, pricing teardowns, and longitudinal studies get picked up because they fill a gap no aggregator can.
Get cited on Reddit, not just by it. Subreddit answers from credible accounts — with a link only when it's actually the best resource — get pulled into Perplexity citations directly. This is community participation, not link building.
Earn a Wikipedia mention the legitimate way. That means being covered in independent, secondary sources first. Wikipedia editors will not accept your own blog as a citation, but they will accept the trade publication that wrote about your data.
Structure for chunk retrieval. Short declarative answers near the top of each section, factual subheads, schema markup, and an llms.txt that points to your most citation-worthy pages. The goal is to make any 200-word slice of your page standalone-useful.
FAQ
Will Reddit's dominance in Perplexity citations decline?
Probably somewhat, as engines diversify sources and as Reddit's content quality on commercial queries gets noisier. But the underlying reason — Reddit is where real questions get real answers — won't disappear. Plan for it to remain a top-three source for years.
Should I create a Wikipedia page for my brand?
Only if you genuinely meet notability guidelines (significant independent coverage in reliable sources). Self-created pages get deleted, and the attempt can backfire. Focus on earning coverage in sources Wikipedia editors already trust.
Do AI engines penalize promotional language directly?
There's no published penalty, but rerankers trained on relevance judgments consistently downweight pages that read as sales copy when the query is informational. The effect is observable across ChatGPT, Perplexity, and Google AI Overviews.


