Field note

Do a few pages earn most of your AI citations? The hero-page pattern

Yes. AI citations tend to pile onto a small set of your pages rather than spread evenly across your site, because engines commit to a tight per-query shortlist and your buyers ask a bounded set of recurring questions. Here's why within-site citation concentration happens, how it differs from cross-web domain concentration, and what to do about your citation hero pages.

Buffy Editorial2026-07-26 · 6 min read

For most sites, a small number of pages earn the overwhelming majority of AI citations, while the rest earn none — the distribution is lopsided, not even. This is a different finding from the well-known one that AI citations concentrate on a few domains across the web. This piece is about the concentration inside your own site: which of the pages you control actually get cited, and why it's so few.

The practical stakes are high. If you spread content effort evenly across your site on the assumption that any page might get cited, you are working against the grain of how engines behave. Knowing which pages are your citation hero pages — and pouring maintenance into those — is one of the highest-leverage moves in AI visibility. Every figure below is attributed, dated, and hedged; the durable claim is the shape of the distribution, not any single percentage.

Why do citations pile onto so few of your pages?

Two mechanisms stack on top of each other. Neither is exotic, and together they make even distribution nearly impossible.

Engines commit to a tight shortlist per answer. An AI answer does not cite every page it reads. It reads broadly and cites narrowly. Perplexity has been reported across 2026 analyses to evaluate roughly ten pages per query but cite only three to five of them, and Google's AI Overviews draw the large majority of their cited sources from the top of the existing organic results rather than deep in the long tail (widely reported in 2026; directional, and it shifts by query type). So for any single question, only a couple of pages win the citation slot.

Your buyers ask a bounded, repeating set of questions. Real buyer queries cluster. The same dozen or two intents — what-is, how-much, versus, best-for, how-to — recur constantly, fanned out into sub-questions. Across that repeating set, the same strong page tends to win the same recurring question again and again.

Put the two together and citations compound onto whichever pages already answer your recurring questions best. A page that wins the definitional query wins it every time it's asked; a page that never wins anything stays at zero. The result is a lopsided curve, not a flat line. This is the same tight-shortlist behaviour that governs cross-web citations, turned inward on your own catalogue.

How is this different from domain concentration?

They sound alike and are constantly confused, so keep the axes separate. One is about the whole web; the other is about your site.

Axis What it measures Typical finding Where we cover it
Cross-web domain concentration Which domains, across the entire web, dominate all AI citations A handful of community and reference domains own a large share; .com/.org alone ~90%+ (Profound, Aug 2024–Jun 2025) Do a few domains dominate AI citations?
Within-site page concentration Which of your own pages get cited when AI answers about you Lopsided: a small set of hero pages earns most; most pages earn none This piece

The domain axis tells you the web is a citation oligarchy and you often need earned placement on sources you don't own. The page axis tells you that even among the pages you do own, effort should not be spread evenly. A brand can rank poorly on the domain axis (rarely cited web-wide) yet still have clear internal hero pages worth defending — and a brand cited web-wide still concentrates onto a few of its own URLs. Both are true at once; neither contradicts the other. Treat them as two dials, measured separately.

What actually makes a page a hero page?

Not seniority or traffic — citability on a recurring question. The pages that become hero pages tend to share traits that make them the easy pick for an engine assembling an answer:

  • They answer one recurring question cleanly and answer-first, so the extractable passage is right there. This is the payoff of answer-first structure.
  • They carry evidence density — a specific attributed statistic, a credible quote, inline sources — which is what lifts a page from read to cited.
  • They are fresh. Citations decay after roughly a quarter, so a page that was a hero six months ago and hasn't been touched can quietly fall off the freshness cliff.
  • They match the format your category rewards. In some verticals the hero is a product or program page; in others it's a comparison or a how-to. That's your citation fingerprint deciding which page type can win at all.

Note what's not on the list: being your homepage, being your newest post, or getting the most human traffic. Hero pages for AI are defined by which of your URLs actually get cited, which is frequently not the page you'd expect.

The mistake isn't publishing too little — it's spreading maintenance evenly across a site when your citations don't. Find the handful of pages engines already reach for, and make those pages undeniable, before you spend a day polishing a page no engine has ever cited.

Why does this matter for how you spend?

Because uniform effort on a lopsided distribution wastes most of the effort. If ten pages earn nearly all your citations and ninety earn none, an even content-maintenance schedule spends 90% of its time where citations don't happen. Reweighting toward the hero pages does three things at once:

  • Defends citations you already have. Hero pages are the ones most exposed to the freshness cliff and to competitors targeting the same recurring question. Letting them go stale is how visibility drops without any obvious cause.
  • Compounds what's working. Adding evidence, updating figures, and building internal links into a page that already gets cited raises the ceiling on a proven winner, rather than gambling on an unproven page.
  • Reveals where you need earned coverage instead. If a recurring, high-value question has no hero page of yours — engines only cite third parties for it — that's the signal to pursue earned placement rather than to keep publishing owned pages the engines ignore.

None of this means stop publishing. New pages are how tomorrow's hero pages are discovered, and depth around a hero page is part of why it wins. It means: publish widely, but maintain narrowly, weighted to the pages the data proves are working. The companion how-to, find your AI citation hero pages, turns this into a repeatable workflow.

What should you do with this?

Treat the shape as durable and measure your own numbers, because the exact ratio is yours alone:

  • Find your hero pages before you assume them. Capture which of your URLs get cited across engines for your buyers' real questions. The winners are often surprising. Keep engines separate — a hero on Perplexity may be invisible on Google AI Mode.
  • Reweight maintenance toward them. Refresh, add evidence, and strengthen internal links on the pages that already earn citations, on a cadence, not once.
  • Flag the gaps. Recurring questions with no owned hero page are earned-media targets, not more-owned-content targets.
  • Re-measure, because the curve moves. Hero pages rise and fall as engines and competitors change. A page list captured once goes stale within a quarter.

The one number that beats every industry average is your own: which of your pages get cited, on which engine, trending which way. Tracking that per-page, per-engine over time — so your hero pages surface and you catch one slipping before your traffic does — is exactly what Buffy Intel is built to do. Questions: [email protected].

Frequently asked

Do most of my AI citations come from a few pages?

For most sites, yes — the distribution is lopsided rather than even. AI engines cite a short shortlist per answer (Perplexity has been reported to evaluate roughly ten pages per query and cite only three to five of them), and your buyers ask a bounded, repeating set of questions. Those two facts push citations onto whichever of your pages already win the recurring questions, so a small number of pages tends to absorb a disproportionate share of your total citations while most of your pages earn none. The exact split varies by site; the shape — concentrated, not uniform — is the durable part. Measure your own before assuming a ratio.

Is this the same as AI citations concentrating on a few domains?

No — it's a different axis. Cross-web domain concentration describes how, across the whole web, a small set of domains (Reddit, Wikipedia, big publishers) dominates all AI citations. Hero-page concentration is within your own site: of the pages you control, only a handful get cited. Both are real and both are 'concentration,' but one is about which domains win the internet and the other is about which of your pages win. A site can have strong internal hero pages and still earn almost no cross-web share, or vice versa.

Should I stop publishing everything except my hero pages?

No. New pages are how you discover future hero pages, and topical depth around a hero page is part of why it gets cited. The point is allocation, not amputation: once you know which pages already earn citations, put your best evidence, freshness, and internal links into reinforcing and expanding them, rather than spreading equal effort across pages that engines never cite. Keep publishing, but weight maintenance toward the pages the data shows are working.