Field note

Do a few domains dominate AI citations? Reconciling the studies

Yes. AI citations concentrate on a relatively small set of domains, and the engines barely overlap with each other. But the exact percentages swing wildly between studies because they measure different things. Here's how to reconcile the conflicting numbers and what the robust finding actually is.

Buffy Editorial2026-07-07 · 6 min read

AI citations concentrate on a relatively small set of domains, and the exact percentages you'll read are all over the map. Both statements are true, and reconciling them is the point of this piece. The concentration is well corroborated: a handful of community and reference sites recur near the top of most engines. The wildly different headline numbers you see quoted are mostly an artefact of studies measuring different things. Different engines, denominators, dates, and methods. Trust the pattern; distrust any single percentage until you know how it was cut.

This is a reconciliation, not a new dataset. It sits alongside our AI search statistics reference and exists to stop a real hazard: pasting one study's number next to another's when the two aren't measuring the same thing. Every figure below is attributed, dated, and hedged.

Do AI citations really concentrate on a few domains?

Yes. The concentration is one of the more robust findings in AI-search research. The clearest quantified cut comes from Profound's "AI Platform Citation Patterns" analysis (Nick Lafferty, published June 2025), covering roughly 680 million citations across ChatGPT, Google AI Overviews, and Perplexity from August 2024 to June 2025:

Finding Figure Source (period)
Share of all citations from .com domains ~80.41% Profound (Aug 2024-Jun 2025)
Share of all citations from .org domains ~11.29% Profound (Aug 2024-Jun 2025)
Wikipedia's share of ChatGPT's top-10 sources ~47.9% Profound (Aug 2024-Jun 2025)
Reddit's share of Perplexity's top-10 sources ~46.7% Profound (Aug 2024-Jun 2025)
Reddit reported as the #1 source across major engines widely reported multiple 2026 analyses

Two TLDs account for over 90% of citations, and within each engine a single source often owns nearly half of the top-10 slots. Profound's own framing is that the concentration is "more extreme than Google PageRank ever produced." Read as a single word, the finding is: concentrated. All figures are vendor-reported and directional.

Why do the studies report such different numbers?

Because the denominator, engine, date, and method are rarely the same, and each changes the number more than the underlying reality does. The cleanest example is Wikipedia inside ChatGPT, where you can find "13%" and "48%" quoted for what sounds like the same thing:

What's being measured Wikipedia figure Source (date)
Wikipedia's share of all US ChatGPT citations ~13.15% 5W Research (June 2026)
Wikipedia's share of ChatGPT's top-10 sources ~47.9% Profound (Aug 2024-Jun 2025)
Wikipedia's share of all ChatGPT citations ~7.8% Profound (Aug 2024-Jun 2025)

These do not contradict each other. The ~48% is a share of a top-10 subset; the ~13% and ~7.8% are shares of all citations, measured by different vendors over different windows. A "share of the top 10" will always dwarf a "share of everything," and a figure from mid-2025 will differ from one a year later. Put a percentage next to another only when the denominator, engine, and date match: otherwise you are comparing a fraction of a shortlist with a fraction of the whole web. This is exactly the kind of answer volatility that makes single snapshots misleading.

How much do the engines overlap with each other?

Barely, and this is the finding that stays stable even as the percentages move. Being cited on one engine tells you little about another. The figures already in our stats reference, from Profound's "State of AI Search" (Zero Click conference, June 2026):

  • Claude ↔ ChatGPT domain overlap: ~8%
  • Claude ↔ Google domain overlap: ~64%
  • Widely reported alongside these: ChatGPT ↔ Perplexity overlap in the low tens of percent. Most domains cited by one are not cited by the other.

Even within one company the surfaces diverge. Google AI Mode and AI Overviews share about 59% of their top-100 sources (BrightEdge, April 2026), yet measured as exact matching URLs rather than shared domains in a top-100 set, reported overlap falls to the low double digits. Same two surfaces, very different numbers, because one counts shared domains and the other counts identical pages. We unpack that pair in do AI Mode and AI Overviews cite the same sources. The lesson generalises: the metric definition drives the number.

Do the engines even cite brands at the same rate?

No, and this is a different metric again, easy to confuse with domain concentration. "How often does an engine cite any brand at all?" is not "which domains does it cite." One 2026 analysis of 34,234 AI responses (Leapd, April 2026) reported brand-citation rates roughly 46× apart between engines:

Engine Reported brand-citation rate Source (date)
ChatGPT ~0.59% Leapd (Apr 2026)
Perplexity ~13.05% Leapd (Apr 2026)
Grok ~27% Leapd (Apr 2026)

A separate analysis (Ranketai, April 2026) put ChatGPT at ~0.7% and Perplexity at ~13.8%. Different absolute numbers, same enormous gap. The corroborated read across both: ChatGPT mentions brands in prose but cites them sparingly, while Perplexity cites densely against a short, high-authority shortlist. This is brand-citation frequency, not source concentration. Keep the two metrics apart when you quote them.

The robust finding across every study is the same shape: citations concentrate on few sources, community content ranks high, engines barely overlap, and the numbers swing by method and month. The specific percentage is the least durable thing in the dataset.

What should you actually do with these numbers?

Use them as directional context, and measure your own reality. Concretely:

  • Always check the cut before you quote. Ask which engine, what denominator (all citations vs a top-N subset), and which date. A number without those three is not comparable to anything.
  • Design for the durable pattern, not the point estimate. Concentration and low cross-engine overlap are stable; the exact shares are not. Build entity strength that travels across engines and earn presence on the community sources that recur near the top. The case for taking Reddit visibility seriously rests on this.
  • Don't let one engine's number stand in for all. With overlap as low as ~8%, a strong showing in ChatGPT says almost nothing about Perplexity. Measure each separately.
  • Refresh, because the numbers decay. Citation shares move month to month, so any figure. Including these. Ages fast. That's the same freshness discipline any competitive reference needs.
  • Treat your own citation coverage as the only number that matters. Whether your brand is cited, on which engine, trending which way, beats every industry average.

The published studies agree on the shape and disagree on the digits, which is exactly why the metric worth watching is your own, tracked per engine over time. That per-engine citation measurement, and how it moves, is what Buffy Intel is built to provide. Questions: [email protected].

Frequently asked

Do AI engines cite the same handful of domains?

They concentrate heavily on a small set of sources, but they don't agree on which set. Community and reference sites. Reddit and Wikipedia especially. Recur near the top of most engines, and one Profound analysis of 680 million citations (August 2024-June 2025) found.com and.org domains alone accounted for roughly 92% of all citations. But cross-engine overlap is low: the same Profound work reported only about 8% domain overlap between Claude and ChatGPT, versus about 64% between Claude and Google. Concentration is real; the specific domains differ by engine.

Why do citation studies report such different percentages for the same source?

Because they measure different things. Wikipedia was reported at about 13.15% of US ChatGPT citations by 5W Research (June 2026) but about 47.9% in another dataset. The second figure is Wikipedia's share of ChatGPT's top-10 sources only, not of all citations, and it covers an earlier period (Profound, August 2024-June 2025). Different denominators, dates, engines, and methods produce different numbers from the same underlying reality. Read the pattern, not the point estimate.

How should I use conflicting AI citation statistics?

Treat published percentages as directional and always check the denominator, engine, and date before quoting one. The durable findings hold across studies: citations concentrate on few sources, community content ranks high, engines barely overlap, and the numbers move month to month. The one number that matters is your own measured citation share on each engine, tracked over time.