Field note

What is information gain, and does it get you cited by AI?

Information gain is how much genuinely new information a page adds beyond what already ranks for the same query. Because AI engines retrieve and cite at the passage level, a page with an original datapoint no competitor has can be cited even when it ranks outside the top results. Here is what the concept means, where it comes from, and why it is becoming the content lever that decides AI citations.

Buffy Editorial2026-08-28 · 7 min read

Information gain is how much genuinely new information a page adds beyond what is already published for the same query. A page that restates the consensus has low information gain; a page with a proprietary statistic, a first-hand test, or a detail no competitor covers has high information gain. It increasingly decides AI citations because engines retrieve and cite at the passage level — a page holding a fact no higher-ranked page provides can be cited even when it ranks outside the top results.

Last reviewed: 28 August 2026. The patent details below are from Google Patents (publication US11354342B2) and contemporaneous coverage (Search Engine Journal). The claim that recent Google updates lean harder on originality is practitioner analysis, attributed and hedged where it appears — Google's own update notes use generic language, so treat the direction as firmer than any single figure.

What is information gain?

Information gain is a score for how much a page adds to what a reader already knows after seeing the pages that already rank. The term entered SEO through a Google patent, "Contextual estimation of link information gain" (publication US11354342B2, filed 2018, granted 2024). The patent describes calculating, for a document, "the additional information" it contains beyond documents a user has previously viewed, then using that score to promote or demote follow-up results.

Aspect Detail
Origin Google patent "Contextual estimation of link information gain" (US11354342B2)
Filed / granted Filed 2018; granted 2024
What it scores Additional information in a document beyond what the user has already seen
Stated use in patent Reranking follow-up results — promoting high-gain, demoting redundant pages
Confirmed in production? No — Google has neither confirmed nor denied it uses this exact mechanism

Source: Google Patents (US11354342B2); patent analysis via Search Engine Journal (2024). The honest reading: information gain is a concept Google has patented and described, not a factor it has confirmed by name. That is enough to treat it as a real lever, and not enough to quote a "score."

Does information gain affect AI citations?

Yes — because AI engines select what to cite at the passage level, not the page level. When a model assembles an answer, it gathers hundreds of candidate content chunks and narrows them to roughly a dozen, keeping the passages that are cleanly extractable, evidence-dense, and non-duplicative. A page whose only content is a restatement of the field adds nothing to that shortlist; a page with a unique, corroborated datapoint earns a slot.

  • Redundancy is a filter, not just a tiebreaker. If five pages say the same thing, an engine needs one of them, not all five. The page with a detail the others lack is the one worth adding to the answer.
  • It partly decouples citation from rank. Practitioner analyses of Google AI Overviews report that the share of citations coming from top-10 organic results fell sharply through 2025–2026 (widely cited as roughly 76% in mid-2025 to about 38% by early 2026 — single-source figures, directional). A page outside the top ten can still be cited if it holds a specific fact no higher-ranked page provides. We cover the rank side of this in does Google rank get you cited by AI and the most-cited domains in AI Overviews.
  • It compounds across the fan-out. A single answer is assembled from many sub-queries. A page with genuinely new material can be the unique source for one branch of the query fan-out, which is how a mid-ranking page gets pulled into an answer it "should not" have made.

If five pages already say it, an engine only needs one of them. Information gain is the reason the sixth page gets cited — it is the one that adds something the other five do not have.

What did Google's March 2026 core update change?

Google confirmed a broad core update that ran 27 March to 8 April 2026, describing it in its usual generic terms as "a regular update designed to better surface relevant, satisfying content for searchers from all types of sites." Google did not name information gain or originality as a target.

What changed is the interpretation. Several practitioners reported that sites leaning on original data and first-hand experience gained visibility while templated, rewritten, and generic AI-generated pages lost it — and framed that as a reweighting toward information gain (for example, DigitalApplied and Evertune analyses, mid-2026; single-vendor, self-reported). Treat those specific magnitudes as unverified. The durable, corroborated pattern underneath is the one the Princeton GEO experiments already established: pages that add statistics, first-hand detail, and cited evidence get lifted; pages that pad word count or restate the field do not. Information gain is the name practitioners now give that pattern, not a new algorithm you can see.

How is information gain different from authority, E-E-A-T, and word count?

Information gain is about the content of the page, not the reputation of the domain or the length of the text. It is easy to confuse with adjacent signals, so separate them:

Signal What it measures Can a small/new site win it?
Information gain New information the page adds vs. what already exists Yes — a unique datapoint beats a big brand's restatement
Domain authority Aggregate link/reputation strength of the domain Rarely quickly — it is slow to build
E-E-A-T Experience, expertise, authoritativeness, trust of the source Partly — first-hand experience helps fast; authority is slow
Content length Word count of the page No — length is not gain; padding lowers signal-to-noise

Source: concept comparison from Google documentation and patent framing. The practical implication is the encouraging one for smaller brands: information gain is the lever least gated by domain authority and page length. A short page with one fact nobody else has can out-cite a long page from a bigger domain that only summarizes the field — which is also why padding a page to hit a word count does not help.

How do you build information gain into a page?

Add something the rest of the field does not have, then make it easy to lift.

  1. Bring original evidence. Data from your own product logs, a first-hand test, a worked example, or an expert observation — anything a competitor cannot copy because they did not do it.
  2. Resolve a contradiction. Where sources disagree, do the reconciliation on the page. A synthesis that settles a real conflict is itself new information.
  3. Cut the restated middle. Remove the boilerplate that repeats what already ranks. High information gain means high signal-to-noise, not more words.
  4. Make the unique fact extractable. State it specifically, numerically, and dated, in a sentence or table a model can lift whole — the chunk-level discipline that gets a passage cited by AI.
  5. Corroborate and attribute. Cite where your claims are data; a unique fact that is also verifiable is safer for an engine to quote.

The test is simple: after reading your page, does a reader — or an engine — know at least one thing they could not get from the pages already ranking? If not, the page is commodity content, and commodity content is the first to fall out of AI answers as it decays.

Information gain is the content half of AI visibility: original, specific, extractable material that gives an engine a reason to cite you rather than the ten pages that all say the same thing. Measuring whether that material actually earns citations — across ChatGPT, Google AI Overviews, Perplexity, and Claude, over time — is exactly what Buffy Intel is built to do. Questions: [email protected].

Frequently asked

What is information gain in SEO?

Information gain is a measure of how much new information a page adds relative to the content already available for the same query. The term comes from a Google patent, 'Contextual estimation of link information gain' (US11354342B2, filed 2018), which describes scoring a document by the additional information it contains beyond documents a user has already seen. A page that restates the existing consensus has low information gain; a page with a proprietary statistic, a first-hand test, or a detail no competitor covers has high information gain. Google has never confirmed it runs this exact scoring in production, so treat it as a well-evidenced concept rather than a named ranking factor.

Does information gain help you get cited by AI?

Yes, indirectly and increasingly. AI engines retrieve and cite at the passage level, narrowing hundreds of candidate chunks to about a dozen and keeping the ones with explicit, verifiable, non-duplicative facts. A page carrying a specific datapoint that no higher-ranked page provides can be pulled into an answer even when it ranks outside the top results — which is why citation has partly decoupled from classic rank. Restated summaries that duplicate the field are the first to be skipped. Information gain is a content-quality lever, not a technical setting.

How do you increase a page's information gain?

Add something the rest of the field does not have: original data from your own logs or tests, first-hand results, a worked example, an expert observation, or a synthesis that resolves a contradiction across existing sources. Cut restated boilerplate, cite where your claims are data, and make the unique fact specific, numeric, and dated so a model can lift it cleanly. The goal is that a reader — or an engine — learns at least one thing on your page they could not get from the pages already ranking.