Field note

Does posting on Reddit get you cited by AI? What the citation data shows

A March–July 2026 citation-mining dataset from Dan Petrovic (DEJAN) found ChatGPT retrieved Reddit 491,024 times but cited it only 3,012 (a 99.4% rejection rate, the highest of any domain), Claude cited Reddit zero times, and only Google cited it heavily. Here is the engine-by-engine data, attributed and hedged, and why Reddit's AI presence tracks its Google ranking.

Buffy Editorial2026-07-24 · 6 min read

Posting on Reddit rarely gets you cited by ChatGPT, effectively never by Claude, and meaningfully only by Google. In a citation-mining dataset covering March–July 2026, published by SEO researcher Dan Petrovic (DEJAN), ChatGPT retrieved Reddit pages 491,024 times but cited only 3,012 of them, a 99.4% rejection rate that Petrovic reports was the highest of any domain in the data. Claude cited Reddit zero times, and Google cited it heavily. Reddit's AI presence, in other words, tracks its ordinary Google ranking, not an AI preference for the platform.

That reframes a year of "put your brand on Reddit to get cited by AI" advice. This piece lays out the engine-by-engine data, attributed and hedged, then reconciles it with what the decision framework for investing in Reddit already says. That piece helps you decide whether to invest; this one shows where the citations actually land across engines.

How often do AI engines actually cite Reddit?

Very differently by engine. Petrovic's method was citation mining: tracking how often each engine retrieves a domain as a candidate versus how often it cites the domain in the final answer, across OpenAI, Google, and Anthropic over roughly six months. The Reddit numbers he reports:

Engine Reddit retrieved Reddit cited Selection rate Reading
ChatGPT (OpenAI) 491,024 3,012 0.61% (99.39% rejected) Retrieved constantly, cited almost never
Claude (Anthropic) Not supplied as a candidate 0 across 139,601 grounding sources ~0% Reddit never even enters the pool
Google Sampled in ~26.6% of searches 14,127 (of 697,768 sources) ~14× ChatGPT's rate Cited heavily

Source: Dan Petrovic / DEJAN, "No, AI doesn't prefer Reddit. Search does." (dejan.ai), citation-mining data March–July 2026; the Anthropic slice covers May–July 2026. Reported secondhand via Search Engine Journal (July 2026). This is single-vendor, self-reported research, so read the exact counts as a directional snapshot of one measurement pipeline, not a fixed constant.

For comparison inside the same OpenAI dataset, Petrovic reports Wikipedia at a 5.64% selection rate and arXiv at 0.77% — so Wikipedia is genuinely favoured on selection, while Reddit and arXiv are retrieved far more than they are kept.

Why does ChatGPT retrieve Reddit so much but cite it so rarely?

Because retrieval and citation are two separate steps, and the gap between them is enormous for Reddit. An engine pulls many candidate pages into context to inform an answer, then footnotes only the few it can read cleanly and that directly answer the question. Being fetched is not being cited; the retrieval-to-citation funnel is documented in detail in how ChatGPT picks the sources it cites, which recorded Reddit fetched roughly 278 times and cited only 11 in a smaller teardown.

Petrovic's data is the same pattern at scale, and it clears up a common confusion: ChatGPT still cites Reddit 3,012 times in absolute terms, which is why other datasets describe ChatGPT as "drawing heavily on Reddit and Wikipedia" (see how many sources each engine cites). Both are true. Reddit's absolute citation count is meaningful, and its selection rate is the worst of any domain measured. The high absolute number comes from sheer retrieval volume, not from the engine trusting each Reddit page. Independent corroboration: Ahrefs' analysis of 1.4 million prompts found Reddit cited in about 1.93% of answers while accounting for a large share of pages ChatGPT pulls in and never names.

Why does Reddit still show up in AI answers, then?

Because a large share of the Reddit citations people actually see come from Google's AI surfaces, and Reddit's presence there is inherited from ordinary Google search ranking. Reddit ranks well in classic Google results, and AI Overviews and AI Mode reuse much of that ranking. Google reported paid data-licensing access to Reddit content, and separately, Reddit is one of the most-cited domains in Google's AI Overviews, second only to YouTube. So the visible Reddit-in-AI phenomenon is real, but it is concentrated on the one engine whose AI layer is fed by web ranking.

Petrovic's one-line thesis captures it:

AI doesn't prefer Reddit. Search does. Reddit's presence in AI answers tracks its organic Google performance, so "post on Reddit to get cited by AI" was always "rank in Google, which happens to feed Google's AI."

Does this mean Reddit "isn't worth it" for AI visibility?

No, and this is where the data must not be over-read. The Reddit decision framework calls Reddit "one of the most-cited sources in AI answers," and that is still accurate on Google's surfaces and in absolute terms. What the new data adds is that the citation is engine-specific: valuable where Google's AI is the destination, weak where ChatGPT or Claude is. That sharpens the framework's "current retrieval share" criterion rather than reversing it.

  • If your buyers' answers come from Google AI Mode / AI Overviews, Reddit remains a live channel, because Google cites it heavily. The play is the ordinary one: earn genuine, helpful mentions that rank in Google search.
  • If your priority is ChatGPT or Claude, Reddit is a weak lever on this data. Effort is better spent on owned, extractable content and earned placement in independent sources those engines actually keep.
  • Never fake it. Seeding threads or buying aged Reddit accounts backfires at both Reddit's moderation layer and the engine's, on every engine.

What are the caveats on this data?

Several, and they matter for how confidently you should act:

  • Single-vendor and self-reported. The counts come from one researcher's measurement pipeline; the specific numbers are a snapshot, not an audited constant. Cite "Petrovic / DEJAN, 2026" with the date when you reuse a figure.
  • Absolute vs rate. Reddit is heavily retrieved and non-trivially cited in absolute terms on ChatGPT; the striking finding is the low selection rate, not a zero.
  • Models change. OpenAI and Anthropic can alter retrieval and grounding at any time; the Anthropic "zero Reddit" finding especially could shift if Claude adds a Reddit source.
  • Inference on Google. Petrovic's Google selection rate is estimated from sampled presence, not a clean retrieved-vs-cited count like OpenAI's, so treat the Google figure as the softest of the three.

None of these undermine the durable, corroborated claim: Reddit is retrieved far more than it is cited, and the citations it does earn are concentrated on Google's search-fed AI.

What should you actually do about Reddit?

Decide per engine, and measure before you commit:

  1. Check your own Reddit citation share by engine. Before assuming Reddit is (or isn't) your channel, look at how often each engine actually cites Reddit for your category's questions. Reddit that ranks in Google for your topic is worth more than Reddit in general.
  2. Match the channel to the destination engine. Google-surface priority favours earned, ranking Reddit presence; ChatGPT/Claude priority favours owned extractable pages and independent third-party placement.
  3. Own your facts; earn your mentions. Publish your specifications, pricing, and definitions as clean, server-rendered text so you are the citable source, and pursue genuine third-party presence for the recommendation, per how to get cited by AI.
  4. Build the durable lever. A corroborated, strong entity across the trusted sources your buyers use outlasts any single-platform tactic.

The through-line is measurement: knowing which engines cite which sources for your buyers' questions, tracked as a trend rather than a one-off snapshot, is exactly what Buffy Intel is built to show — engine by engine, over time.

Frequently asked

Does posting on Reddit get you cited by AI engines?

Rarely on ChatGPT, effectively never on Claude, and meaningfully only on Google. In a March–July 2026 citation-mining dataset published by Dan Petrovic (DEJAN), ChatGPT retrieved Reddit pages 491,024 times but cited only 3,012 of them, a 99.4% rejection rate that was the highest of any domain measured. Claude cited Reddit zero times across 139,601 grounding sources and did not even pull it in as a candidate. Google cited Reddit heavily. It is single-vendor, directional data, so treat the exact figures as a snapshot, but the engine split is the point: Reddit is a Google-surface play, not a ChatGPT or Claude one.

Why does Reddit still appear in so many AI answers if ChatGPT rejects it?

Because a large share of visible Reddit citations come from Google's AI surfaces, and Reddit's presence there tracks its ordinary Google search ranking rather than any AI-specific preference. Reddit ranks well in classic Google results, and Google's AI Overviews and AI Mode inherit those rankings. Petrovic's framing is that 'AI doesn't prefer Reddit, search does.' So 'post on Reddit to get cited by AI' has largely meant 'rank in Google, which happens to feed Google's AI.'

Should I stop investing in Reddit for AI visibility?

Not necessarily, but decide it per engine. If your buyers and AI answers about your category lean on Google's AI surfaces, Reddit can still earn you visibility, because Google cites it heavily. If your priority is ChatGPT or Claude citations, the data suggests Reddit is a weak lever, and owned, extractable content plus earned placement in independent sources will do more. Measure your own Reddit citation share by engine before committing a team to the channel.