Field note

Which content formats survive in AI citations? The 'content graveyard' data

A 2026 study of 2.4M citations across 28,725 domains found 57.2% were cited in a single month and never again — while only 2.7% held a citation across all seven months. The differentiator wasn't brand size, it was format: comparison, FAQ, how-to and pricing pages persisted at roughly twice the rate, and 'complete guide' pages were 3.5x more common among the domains that vanished. Here's the fully-attributed data and what it means for what you publish.

Buffy Editorial2026-08-18 · 7 min read

In a 2026 study of 2.4 million AI citations, only 2.7% of domains held a citation across all seven months, while 57.2% were cited in a single month and then never again — and the thing that separated the survivors was content format, not brand size. Somantra AI's "Content Graveyard" analysis found that comparison, FAQ, how-to and pricing pages persisted at roughly twice the rate, while pages labelled a "complete guide" were about 3.5x more common among the domains that vanished after one citation.

Last reviewed: 18 August 2026. All figures below come from Somantra AI's Content Graveyard analysis (published 17 August 2026), a single-vendor read of one sector (Australian insurance) across ChatGPT and Google over seven months. It is directional, not definitive — cite "Somantra AI, 2026" with the date when you reuse a number, and treat the pattern (format predicts durability) as firmer than any single multiple.

What did the 'content graveyard' study find?

That most of the web an AI engine cites is cited once and then forgotten. Somantra AI tracked 2,437,107 citation records across 28,725 domains on ChatGPT and Google AI Overviews over a seven-month window (reported as November 2025 to July 2026) in the Australian insurance sector. The headline split is stark.

Metric Value
Citation records analysed 2,437,107
Distinct domains tracked 28,725
Engines ChatGPT, Google (AI Overviews)
Window 7 months (Nov 2025 – Jul 2026, reported)
Cited in exactly one month, never again 57.2% of domains
Cited in all seven months 2.7% of domains
Sector Australian insurance

Source: Somantra AI, Content Graveyard, 2026. The one-line read: being cited once is common and cheap; staying cited is rare. Because it is one vendor measuring one vertical, read the magnitudes as directional — but a >20:1 ratio between one-and-done and always-present domains is a large effect by any measure, and it reframes the goal from getting cited to staying cited.

Which formats survived, and which vanished?

The gap between the "content graveyard" and the durable minority was separated by format, not by how big the brand was. Structured, comparative and specific pages persisted; generic long-form guides did not.

Content format Citation behaviour
Comparison-table content Among long-term survivors at ~2x the rate of one-citation domains
Pricing / discount / savings pages ~2x more common among long-term survivors
FAQ and how-to pages Correlated with persistence across the seven-month window
"Complete guide" pages ~3.5x more common among domains that vanished after a single citation

Source: Somantra AI, 2026. The pattern is consistent: the formats that survived are the ones that hand a model a clean, self-contained, specific answer — a row in a comparison, a priced option, a direct question-and-answer. The format that vanished is the one that buries the answer inside sprawling prose. As the study's authors put it, the differentiator is reproducible precisely because it is not about brand size:

"If persistence tracked brand size, a smaller insurer would have no route into a durable citation position. Because it tracks format, the formula is reproducible."

Does this mean long-form content is dead?

No — and this is where the finding is easy to misread. A separate 2026 analysis of AI-Overview citations found that word count barely predicts whether a page is cited at all; the real lever is coverage, because comprehensive pages answer more of the sub-questions a query fans out into, as we covered in does content length affect AI citations. Length and citation capture travel together only because thorough pages cover more branches.

The "content graveyard" data is about a different axis: durability, not capture. A long "complete guide" can win a citation on breadth this month and lose it next month, because a generic guide is interchangeable — a dozen other guides say the same thing, so the engine has no reason to keep choosing yours. A comparison table with your specific numbers, or an FAQ that answers one exact question, is harder to swap out. So the lesson is not "write less." It is: make each section a specific, self-contained, comparative answer — the extractable chunk a model keeps selecting — rather than one long undifferentiated read.

Why would format predict citation survival?

Because retrieval and citation happen at the passage level, and structured formats win the passage-level contest repeatedly. Three mechanisms line up.

  1. Extractability. A model gathering candidate passages keeps the ones with clean boundaries — a table row, a priced line item, a Q-and-A pair — over a point buried mid-paragraph. Structured formats are pre-chunked; a "complete guide" makes the engine do the chunking, and an easier source usually wins.
  2. Specificity. Comparison and pricing pages carry named, numeric, dated facts. Those are exactly the claims engines lift and the ones least likely to be duplicated by a competitor's generic page, so they hold their position longer.
  3. Reproducible substitution. When the cited pool churns, the engine replaces a source with the next-best answer to that sub-query. A generic guide has many near-identical substitutes; a specific comparison or spec table has fewer, so it survives more rounds of substitution.

Note these are correlational — Somantra observed format travelling with persistence, not a controlled experiment isolating format as the cause. But the mechanism is the same one that governs how answer engines select any chunk, which is why the direction is credible.

Does this contradict our citation-turnover and freshness pieces?

No — it completes them. We have written that the cited pool churns fast (only about 10.6% of URLs survived six weeks in how long AI citations last) and that live retrieval favours recently-updated pages. Those pieces answer how much the pool turns over and how fast a page decays. The "content graveyard" data answers a third question: given that the pool churns, who survives it? The reconciliation is clean:

  • Turnover says the set of cited sources is mostly replaced within weeks — true across studies.
  • Freshness says stale pages get displaced first — a timing lever you control by updating.
  • Format durability says that among pages competing to be re-cited, structured and specific formats win more rounds — a structural lever you control by how you write.

None of these say a citation is permanent. Together they say the same thing from three angles: a citation is a position you defend, and the defensible positions are fresh, specific, and structured. There is no contradiction — freshness is the clock, format is the moat.

What should you actually do about it?

Publish for durability, not just for the first citation. The practical moves follow directly from the data:

  1. Vary the format deliberately. Don't default every page to a long explainer. Mix in comparison tables, FAQ blocks, how-to steps, and pages that carry specific pricing or spec data — the formats that persisted.
  2. Pre-chunk your own pages. Inside any long guide, break the answer into self-contained, question-led sections with tables and lists, so the engine lifts a clean unit instead of skipping the page.
  3. Be specific where competitors are generic. Named, numeric, dated facts are both more citable and more durable, because they are harder to substitute.
  4. Watch durability, not just presence. A single citation snapshot tells you nothing about whether you'll be there next month. Track share of voice and citation coverage as trends across many prompts and engines, and refresh the pages that matter before they age out.

The mistake the "content graveyard" exposes is treating a first citation as the finish line. On this data, most domains reach that line once and never return — and the ones that stay did it by publishing formats a model could keep choosing.


Buffy Intel tracks which of your pages are cited across AI engines, and — just as importantly — whether they stay cited over time, so you can see the difference between a one-month appearance and a durable position. If you want to know which of your formats are surviving the churn and which are already in the graveyard, start with Buffy Intel or reach us at [email protected].

Frequently asked

Which content formats last longest in AI citations?

On the best format-level read available, structured and comparative formats last longest. Somantra AI's 2026 'Content Graveyard' analysis of 2.4 million citations found that comparison-table content, FAQ and how-to pages, and pricing or discount pages appeared among long-term survivors at roughly twice the rate of one-citation domains. Pages labelled a 'complete guide' were the opposite: about 3.5 times more common among domains that were cited once and never again. It is single-vendor, single-sector data, so read it as directional, but it lines up with how retrieval works — engines lift clean, self-contained chunks, and structured formats hand them one.

Does this mean long-form guides are bad for AI search?

No. A separate 2026 analysis found that comprehensive pages get cited more often because they cover more of the sub-questions a query fans out into — the lever is coverage, not word count. The 'content graveyard' finding is about a different thing: durability, or whether a citation is held over months. A sprawling 'complete guide' can win a citation on breadth and then lose it, because it is generic and easily replaced. The fix is not to write shorter; it is to make each section a self-contained, comparative, specific answer a model can keep choosing.

Is one study enough to change what I publish?

Treat it as a strong signal, not a law. The 'content graveyard' data is single-vendor (Somantra AI), single-sector (Australian insurance), covers two engines (ChatGPT and Google), and is correlational — it shows format travels with persistence, not that format alone causes it. But it points the same way as the wider evidence that AI engines over-index on a few extractable formats, and it costs nothing to act on: mixing comparison tables, FAQs and specific pricing or spec data into your pages is good practice whether or not the exact multiples hold.