Information gain is how much genuinely new information a page adds beyond what is already published for the same query. A page that restates the consensus has low information gain; a page with a proprietary statistic, a first-hand test, or a detail no competitor covers has high information gain. It increasingly decides AI citations because engines retrieve and cite at the passage level — a page holding a fact no higher-ranked page provides can be cited even when it ranks outside the top results.
Last reviewed: 28 August 2026. The patent details below are from Google Patents (publication US11354342B2) and contemporaneous coverage (Search Engine Journal). The claim that recent Google updates lean harder on originality is practitioner analysis, attributed and hedged where it appears — Google's own update notes use generic language, so treat the direction as firmer than any single figure.
What is information gain?
Information gain is a score for how much a page adds to what a reader already knows after seeing the pages that already rank. The term entered SEO through a Google patent, "Contextual estimation of link information gain" (publication US11354342B2, filed 2018, granted 2024). The patent describes calculating, for a document, "the additional information" it contains beyond documents a user has previously viewed, then using that score to promote or demote follow-up results.
| Aspect | Detail |
|---|---|
| Origin | Google patent "Contextual estimation of link information gain" (US11354342B2) |
| Filed / granted | Filed 2018; granted 2024 |
| What it scores | Additional information in a document beyond what the user has already seen |
| Stated use in patent | Reranking follow-up results — promoting high-gain, demoting redundant pages |
| Confirmed in production? | No — Google has neither confirmed nor denied it uses this exact mechanism |
Source: Google Patents (US11354342B2); patent analysis via Search Engine Journal (2024). The honest reading: information gain is a concept Google has patented and described, not a factor it has confirmed by name. That is enough to treat it as a real lever, and not enough to quote a "score."
Does information gain affect AI citations?
Yes — because AI engines select what to cite at the passage level, not the page level. When a model assembles an answer, it gathers hundreds of candidate content chunks and narrows them to roughly a dozen, keeping the passages that are cleanly extractable, evidence-dense, and non-duplicative. A page whose only content is a restatement of the field adds nothing to that shortlist; a page with a unique, corroborated datapoint earns a slot.
- Redundancy is a filter, not just a tiebreaker. If five pages say the same thing, an engine needs one of them, not all five. The page with a detail the others lack is the one worth adding to the answer.
- It partly decouples citation from rank. Practitioner analyses of Google AI Overviews report that the share of citations coming from top-10 organic results fell sharply through 2025–2026 (widely cited as roughly 76% in mid-2025 to about 38% by early 2026 — single-source figures, directional). A page outside the top ten can still be cited if it holds a specific fact no higher-ranked page provides. We cover the rank side of this in does Google rank get you cited by AI and the most-cited domains in AI Overviews.
- It compounds across the fan-out. A single answer is assembled from many sub-queries. A page with genuinely new material can be the unique source for one branch of the query fan-out, which is how a mid-ranking page gets pulled into an answer it "should not" have made.
If five pages already say it, an engine only needs one of them. Information gain is the reason the sixth page gets cited — it is the one that adds something the other five do not have.
What did Google's March 2026 core update change?
Google confirmed a broad core update that ran 27 March to 8 April 2026, describing it in its usual generic terms as "a regular update designed to better surface relevant, satisfying content for searchers from all types of sites." Google did not name information gain or originality as a target.
What changed is the interpretation. Several practitioners reported that sites leaning on original data and first-hand experience gained visibility while templated, rewritten, and generic AI-generated pages lost it — and framed that as a reweighting toward information gain (for example, DigitalApplied and Evertune analyses, mid-2026; single-vendor, self-reported). Treat those specific magnitudes as unverified. The durable, corroborated pattern underneath is the one the Princeton GEO experiments already established: pages that add statistics, first-hand detail, and cited evidence get lifted; pages that pad word count or restate the field do not. Information gain is the name practitioners now give that pattern, not a new algorithm you can see.
How is information gain different from authority, E-E-A-T, and word count?
Information gain is about the content of the page, not the reputation of the domain or the length of the text. It is easy to confuse with adjacent signals, so separate them:
| Signal | What it measures | Can a small/new site win it? |
|---|---|---|
| Information gain | New information the page adds vs. what already exists | Yes — a unique datapoint beats a big brand's restatement |
| Domain authority | Aggregate link/reputation strength of the domain | Rarely quickly — it is slow to build |
| E-E-A-T | Experience, expertise, authoritativeness, trust of the source | Partly — first-hand experience helps fast; authority is slow |
| Content length | Word count of the page | No — length is not gain; padding lowers signal-to-noise |
Source: concept comparison from Google documentation and patent framing. The practical implication is the encouraging one for smaller brands: information gain is the lever least gated by domain authority and page length. A short page with one fact nobody else has can out-cite a long page from a bigger domain that only summarizes the field — which is also why padding a page to hit a word count does not help.
How do you build information gain into a page?
Add something the rest of the field does not have, then make it easy to lift.
- Bring original evidence. Data from your own product logs, a first-hand test, a worked example, or an expert observation — anything a competitor cannot copy because they did not do it.
- Resolve a contradiction. Where sources disagree, do the reconciliation on the page. A synthesis that settles a real conflict is itself new information.
- Cut the restated middle. Remove the boilerplate that repeats what already ranks. High information gain means high signal-to-noise, not more words.
- Make the unique fact extractable. State it specifically, numerically, and dated, in a sentence or table a model can lift whole — the chunk-level discipline that gets a passage cited by AI.
- Corroborate and attribute. Cite where your claims are data; a unique fact that is also verifiable is safer for an engine to quote.
The test is simple: after reading your page, does a reader — or an engine — know at least one thing they could not get from the pages already ranking? If not, the page is commodity content, and commodity content is the first to fall out of AI answers as it decays.
Information gain is the content half of AI visibility: original, specific, extractable material that gives an engine a reason to cite you rather than the ten pages that all say the same thing. Measuring whether that material actually earns citations — across ChatGPT, Google AI Overviews, Perplexity, and Claude, over time — is exactly what Buffy Intel is built to do. Questions: [email protected].