In a 2026 study of 2.4 million AI citations, only 2.7% of domains held a citation across all seven months, while 57.2% were cited in a single month and then never again — and the thing that separated the survivors was content format, not brand size. Somantra AI's "Content Graveyard" analysis found that comparison, FAQ, how-to and pricing pages persisted at roughly twice the rate, while pages labelled a "complete guide" were about 3.5x more common among the domains that vanished after one citation.
Last reviewed: 18 August 2026. All figures below come from Somantra AI's Content Graveyard analysis (published 17 August 2026), a single-vendor read of one sector (Australian insurance) across ChatGPT and Google over seven months. It is directional, not definitive — cite "Somantra AI, 2026" with the date when you reuse a number, and treat the pattern (format predicts durability) as firmer than any single multiple.
What did the 'content graveyard' study find?
That most of the web an AI engine cites is cited once and then forgotten. Somantra AI tracked 2,437,107 citation records across 28,725 domains on ChatGPT and Google AI Overviews over a seven-month window (reported as November 2025 to July 2026) in the Australian insurance sector. The headline split is stark.
| Metric | Value |
|---|---|
| Citation records analysed | 2,437,107 |
| Distinct domains tracked | 28,725 |
| Engines | ChatGPT, Google (AI Overviews) |
| Window | 7 months (Nov 2025 – Jul 2026, reported) |
| Cited in exactly one month, never again | 57.2% of domains |
| Cited in all seven months | 2.7% of domains |
| Sector | Australian insurance |
Source: Somantra AI, Content Graveyard, 2026. The one-line read: being cited once is common and cheap; staying cited is rare. Because it is one vendor measuring one vertical, read the magnitudes as directional — but a >20:1 ratio between one-and-done and always-present domains is a large effect by any measure, and it reframes the goal from getting cited to staying cited.
Which formats survived, and which vanished?
The gap between the "content graveyard" and the durable minority was separated by format, not by how big the brand was. Structured, comparative and specific pages persisted; generic long-form guides did not.
| Content format | Citation behaviour |
|---|---|
| Comparison-table content | Among long-term survivors at ~2x the rate of one-citation domains |
| Pricing / discount / savings pages | ~2x more common among long-term survivors |
| FAQ and how-to pages | Correlated with persistence across the seven-month window |
| "Complete guide" pages | ~3.5x more common among domains that vanished after a single citation |
Source: Somantra AI, 2026. The pattern is consistent: the formats that survived are the ones that hand a model a clean, self-contained, specific answer — a row in a comparison, a priced option, a direct question-and-answer. The format that vanished is the one that buries the answer inside sprawling prose. As the study's authors put it, the differentiator is reproducible precisely because it is not about brand size:
"If persistence tracked brand size, a smaller insurer would have no route into a durable citation position. Because it tracks format, the formula is reproducible."
Does this mean long-form content is dead?
No — and this is where the finding is easy to misread. A separate 2026 analysis of AI-Overview citations found that word count barely predicts whether a page is cited at all; the real lever is coverage, because comprehensive pages answer more of the sub-questions a query fans out into, as we covered in does content length affect AI citations. Length and citation capture travel together only because thorough pages cover more branches.
The "content graveyard" data is about a different axis: durability, not capture. A long "complete guide" can win a citation on breadth this month and lose it next month, because a generic guide is interchangeable — a dozen other guides say the same thing, so the engine has no reason to keep choosing yours. A comparison table with your specific numbers, or an FAQ that answers one exact question, is harder to swap out. So the lesson is not "write less." It is: make each section a specific, self-contained, comparative answer — the extractable chunk a model keeps selecting — rather than one long undifferentiated read.
Why would format predict citation survival?
Because retrieval and citation happen at the passage level, and structured formats win the passage-level contest repeatedly. Three mechanisms line up.
- Extractability. A model gathering candidate passages keeps the ones with clean boundaries — a table row, a priced line item, a Q-and-A pair — over a point buried mid-paragraph. Structured formats are pre-chunked; a "complete guide" makes the engine do the chunking, and an easier source usually wins.
- Specificity. Comparison and pricing pages carry named, numeric, dated facts. Those are exactly the claims engines lift and the ones least likely to be duplicated by a competitor's generic page, so they hold their position longer.
- Reproducible substitution. When the cited pool churns, the engine replaces a source with the next-best answer to that sub-query. A generic guide has many near-identical substitutes; a specific comparison or spec table has fewer, so it survives more rounds of substitution.
Note these are correlational — Somantra observed format travelling with persistence, not a controlled experiment isolating format as the cause. But the mechanism is the same one that governs how answer engines select any chunk, which is why the direction is credible.
Does this contradict our citation-turnover and freshness pieces?
No — it completes them. We have written that the cited pool churns fast (only about 10.6% of URLs survived six weeks in how long AI citations last) and that live retrieval favours recently-updated pages. Those pieces answer how much the pool turns over and how fast a page decays. The "content graveyard" data answers a third question: given that the pool churns, who survives it? The reconciliation is clean:
- Turnover says the set of cited sources is mostly replaced within weeks — true across studies.
- Freshness says stale pages get displaced first — a timing lever you control by updating.
- Format durability says that among pages competing to be re-cited, structured and specific formats win more rounds — a structural lever you control by how you write.
None of these say a citation is permanent. Together they say the same thing from three angles: a citation is a position you defend, and the defensible positions are fresh, specific, and structured. There is no contradiction — freshness is the clock, format is the moat.
What should you actually do about it?
Publish for durability, not just for the first citation. The practical moves follow directly from the data:
- Vary the format deliberately. Don't default every page to a long explainer. Mix in comparison tables, FAQ blocks, how-to steps, and pages that carry specific pricing or spec data — the formats that persisted.
- Pre-chunk your own pages. Inside any long guide, break the answer into self-contained, question-led sections with tables and lists, so the engine lifts a clean unit instead of skipping the page.
- Be specific where competitors are generic. Named, numeric, dated facts are both more citable and more durable, because they are harder to substitute.
- Watch durability, not just presence. A single citation snapshot tells you nothing about whether you'll be there next month. Track share of voice and citation coverage as trends across many prompts and engines, and refresh the pages that matter before they age out.
The mistake the "content graveyard" exposes is treating a first citation as the finish line. On this data, most domains reach that line once and never return — and the ones that stay did it by publishing formats a model could keep choosing.
Buffy Intel tracks which of your pages are cited across AI engines, and — just as importantly — whether they stay cited over time, so you can see the difference between a one-month appearance and a durable position. If you want to know which of your formats are surviving the churn and which are already in the graveyard, start with Buffy Intel or reach us at [email protected].