Field note

Where do AI citations come from? Earned media, owned pages, and paid content

Overwhelmingly from earned media — independent third-party coverage, not brand-owned or paid pages. In Muck Rack's May 2026 analysis of over 25 million links cited by ChatGPT, Claude, and Gemini, earned media drove about 84% of citations and paid or advertorial content just 0.3%. Here is the source-type breakdown, the per-engine differences, and what it means for where you spend your GEO effort.

Buffy Editorial2026-08-25 · 5 min read

Overwhelmingly from earned media — independent third-party coverage such as news, reviews, and forums, not the pages a brand publishes or pays for. In Muck Rack's May 2026 "What Is AI Reading?" analysis of over 25 million links cited by ChatGPT, Claude, and Gemini, earned media drove about 84% of citations, while paid or advertorial content accounted for just 0.3%. This reference lays out the source-type breakdown, how it differs by engine, and what it means for where you spend your GEO effort.

Last reviewed: 25 August 2026. All figures below are from Muck Rack's "What Is AI Reading?" study (May 2026 edition, the third since July 2025) unless noted. It is a single-vendor analysis across three engines, so treat the direction as firm and any single percentage as directional, and cite "Muck Rack, 2026" with the date when you reuse a figure.

Where do AI engines get the sources they cite?

Mostly from earned media, by a wide margin, with paid content almost nowhere. Here is the headline split across the 25 million-plus cited links Muck Rack analysed:

Source type Share of AI citations Note
Earned media (total) ~84% Independent third-party coverage
Journalism (subset of earned) ~27% Across 20,000+ distinct outlets
Paid / advertorial ~0.3% Effectively negligible

Source: Muck Rack "What Is AI Reading?", May 2026. The one-line summary: AI engines assemble answers out of what independent third parties have said, not out of what brands say about themselves or pay to place. That is the same pattern behind why independent best-of lists get cited far more than brand-owned pages for commercial "best X" queries.

Does the earned-media dominance hold across engines and over time?

Yes on both counts, which is what makes it worth planning around. Across three editions of the study since July 2025, earned media's share stayed in an 82–89% band and journalism in a 25–27% band — a stable pattern, not a single snapshot. But which sources each engine reaches for varies sharply:

Engine Cites sources in… Avg. citations when it cites Top cited domain
ChatGPT ~96% of responses ~5 Wikipedia
Gemini ~82% of responses ~8 Reddit
Claude ~55% of responses ~13 PubMed Central

Source: Muck Rack, 2026. Note this counts cited links shown per response on Muck Rack's corpus, which is a different measurement from the sources pulled per answer reported elsewhere — Semrush's 2026 index put ChatGPT at about 15 sources and Gemini at about 3 using its own method. Read the two as complementary cuts, not conflicting numbers: the metrics, corpora, and dates differ, so do not stack them. The durable takeaway is the shape — the same earned-media dominance flows through a different favourite source on each engine, and one engine's staple (Reddit for Gemini, PubMed Central for Claude) can be near-absent on another. One more per-engine marker: Muck Rack found Axios in ChatGPT's top three cited domains across 13 of 17 industries.

Why do AI engines lean on earned media instead of your own pages?

Because independent coverage is the corroboration signal engines trust most. A claim a brand makes about itself is one voice; the same claim carried by journalists, reviewers, and forums is many independent voices, which is exactly the evidence an answer engine weights when it decides whom to cite. Muck Rack's own read is consistent across editions: earned media is the dominant input, and it also found press releases appearing about 3.5 times more often in industry-trend answers than in "best-of" answers, so the type of earned coverage that lands depends on the question being asked.

This is why earned media is the durable lever and why faking it backfires — manufactured coverage reads as manipulation and can trigger penalties rather than citations.

Does this mean your own pages don't matter?

No — it means owned and earned pages win different jobs. Muck Rack's ~84%-earned figure sits alongside a second finding from its 2026 State of PR survey that roughly 99% of AI citations come from non-paid sources. The two reconcile cleanly: paid is ~0.3%, so nearly everything is non-paid, and within that non-paid pool earned media is the ~84% majority and brand-owned pages are the remainder. Owned pages still earn the first-party citations that matter for brand- and feature-specific questions, where your own site is the authoritative source.

Earned media is how AI answers decide who to trust; owned pages are how they get the facts right. You need both, and paid placement earns you almost neither.

There is a live wrinkle worth dating: on ChatGPT specifically, a mid-2026 shift toward fetching first-party domains directly after GPT-5.6 is nudging more of its citations toward brand-owned pages. That is an engine-specific, post-May-2026 movement, and it does not overturn the cross-engine, through-mid-2026 earned-media dominance measured here — it is a trend to watch alongside it, not a contradiction of it.

What does this mean for your GEO strategy?

Spend where the citations actually come from: earned first, owned for the facts, paid almost never.

  • Pursue earned placement for evaluative and "best X" questions. Independent reviews, journalism, and community coverage are where ~84% of citations originate — see how to get into AI-cited best lists and how to use Reddit for AI-search visibility.
  • Keep owned pages clean and factual for brand and feature queries. Server-rendered, extractable specs and pricing are what an engine lifts when it does cite your domain — the discipline in prioritising your structured data.
  • Do not buy your way in. At ~0.3% of citations, paid and advertorial content is the weakest AI-visibility lever there is.
  • Measure per engine. Because each engine leans on a different top source, track ChatGPT, Gemini, and Claude separately rather than trusting one blended number — the reasoning in who owns AI search visibility.

The honest read for late 2026: the source of an AI citation is far more often something a third party wrote about you than anything you published or paid for — so earned credibility, not self-promotion, is the core of AI visibility. Watching which earned and owned sources actually get cited for your brand, per engine and over time, is exactly what Buffy Intel measures. Questions: [email protected].

Frequently asked

What share of AI citations come from earned media?

About 84%, per Muck Rack's May 2026 'What Is AI Reading?' study of more than 25 million links cited by ChatGPT, Claude, and Gemini. Journalism alone accounted for about 27% of cited sources, spanning more than 20,000 distinct outlets, while paid and advertorial content made up just 0.3% of citations. The earned-media share has held in an 82–89% band across three editions of the study since July 2025, so read it as a stable, directional pattern rather than a one-off reading. It is single-vendor data — attribute 'Muck Rack, 2026' and hedge the exact figure.

Do AI engines cite paid or advertorial content?

Almost never. Muck Rack's 2026 analysis put paid and advertorial content at about 0.3% of AI citations — effectively a rounding error next to earned media's ~84%. AI engines lean on independent, corroborated sources when they decide what to cite, so buying placements is one of the weakest ways to earn an AI citation. The durable route is genuine third-party coverage and clean, extractable facts on your own pages.

Does each AI engine cite sources differently?

Yes, sharply. In Muck Rack's 2026 data, ChatGPT cited sources in about 96% of responses (top domain Wikipedia), Gemini in about 82% (top domain Reddit), and Claude in about 55% but with more citations when it did cite (top domain PubMed Central). So the same earned-media dominance plays out through different favourite sources per engine. Measure each engine separately, because a source that carries weight in one may be near-invisible in another.