An information-gain audit finds the pages that only restate what already ranks — the ones AI engines read and skip — and shows you where to add the original data, first-hand detail, or unique synthesis that earns a citation. The method below is a repeatable, five-step review you can run without a special tool: score each page for what it adds, flag the commodity content, and fix the highest-value gaps first.
Last reviewed: 28 August 2026. This is a method built on how AI engines select passages (chunk-level retrieval that filters redundant content) and on the information gain concept from Google's patent literature. It reconciles our existing guidance on what content changes lift AI citations into an audit you can run on your own corpus.
How do you audit a page for information gain?
Compare the page against the pages already answering the query, and measure the gap. Work one target query per page, in five steps:
- Run the query and read the field. Search the query your page targets and open the pages currently ranking or being cited in AI answers. These are what an engine already has.
- List the shared consensus. Write down every fact, claim, and number that appears across those pages. This is the information an engine can already assemble without you.
- Mark what your page adds. Read your page and highlight only what is not on the shared list — an original datapoint, a first-hand result, a worked example, a detail no competitor covers.
- Score the gap. If step 3 is empty, the page has low information gain (commodity content). If it holds one or more unique, verifiable facts, it has real gain. Rank your pages by that gap.
- Fix highest-value first. Prioritise pages where you have original evidence to add and that already get crawled but not cited — the fastest commodity-to-citable moves.
The output is a list of pages sorted by how much they add, not by traffic or length. That sort is the audit.
Which pages have low information gain?
The commodity pages share a fingerprint. Use this table to flag them fast during the audit:
| Signal | What it looks like | Why it fails |
|---|---|---|
| Restated consensus | Every claim also appears on the top competitors | Adds nothing to the answer shortlist |
| No original evidence | No first-hand data, tests, or examples | Nothing unique for a model to lift |
| Padded length | Long, but high word-to-fact ratio | Low signal-to-noise; length is not gain |
| Rewritten source | A paraphrase of one or two ranking pages | Duplicate of what the engine already has |
| Crawled, not cited | AI bots hit it; citations stay flat | Read and passed over as redundant |
A page matching two or more rows is commodity content — it duplicates the field and will keep being skipped. The last row is the leading indicator: heavy crawler activity with no citations often means the content is being evaluated and rejected for redundancy, not for reachability.
The fastest way to find your low-gain pages is to ask, for each one: if this page vanished, would the answer to its query be any worse? If not, an engine already agrees — that is why it is not citing you.
How do you add information gain to a thin page?
Give the page something the field does not have, then make it liftable. In priority order:
- Add original data. A number from your own logs, a test result, or a small study is the highest-value gain because no competitor can copy it.
- Add first-hand detail. A worked example, a screenshot-backed walkthrough, or an expert observation from real use — experience the ranking pages lack.
- Resolve a contradiction. Where sources disagree, do the reconciliation on the page; a synthesis that settles a real conflict is new information in itself.
- Cut the restated middle. Remove boilerplate that repeats the consensus so the unique material stands out and signal-to-noise rises.
- Make the unique fact extractable. State it specifically, numerically, and dated in a sentence or table a model can lift whole — the extractable-chunk discipline — and corroborate it so it is safe to quote.
Re-score the page after editing: read the field again, and confirm your page now holds at least one fact the ranking pages do not. If it does, you have moved it from commodity to citable.
How often should you re-run the audit?
Put competitive pages on a refresh cadence, because information gain erodes as the rest of the field catches up. A datapoint that was unique when you published it becomes consensus once competitors copy it, and a page that once stood out slides back to commodity. That is the same content-decay cliff that pulls stale pages out of AI answers — from the content side rather than the freshness side.
- Competitive, fast-moving pages: re-audit each quarter; the field moves and your gain shrinks.
- Evergreen or definitional pages: re-audit rarely — do not churn stable content just to restate it.
- After any ranking or citation drop: audit the affected pages for gain before assuming a technical cause, and fix the content the durable way rather than spam-proofing a healthy page.
An information-gain audit turns "write better content" into a concrete, repeatable pass: find the pages that add nothing, add something real, and make it liftable. Tracking whether those edits actually earn citations — across ChatGPT, Google AI Overviews, Perplexity, and Claude, over time — is exactly what Buffy Intel is built to do. Questions: [email protected].