Information Gain
What information gain means in SEO and AI search — how Google's Information Gain Score patent works, why original research and unique data outperform regurgitated content, and how to measure it.
Information gain is how much *new* information your page adds beyond what a searcher has already seen in the results for a topic — novelty relative to the existing corpus, not general content quality. The term comes from a real Google patent ('Contextual estimation of link information gain,' filed 2018, granted 2022) that scores documents 0.00–1.00 for novelty and ranks partly on that score — but Google has never confirmed using it in live ranking, and the patent's framing is about what to show *next* after a user has seen some results, not the first page. There's no visible score in Search Console, no API, and no public formula: any 'information gain score' in a tool is an approximation. What's real and official is Google's helpful-content guidance asking whether content offers 'original information, reporting, research, or analysis.' It matters more now because AI Overviews and AI Mode can already synthesize consensus content — only genuinely new information (proprietary data, original research, first-hand expertise) is what AI has to cite instead of absorb.
TL;DR — Information gainInformation gain is how much new information a page adds beyond what a searcher has already seen in prior results on the same topic — novelty relative to the existing corpus, not general content quality. The term comes from a granted Google patent that Google has never confirmed using in live ranking. is how much new information your page adds that a searcher hasn’t already seen in the other results for that topic. It’s not the same as “good content” or “long content” — a page can be thorough and still add nothing new. The name comes from a Google patent, but Google has never said it actually uses it to rank pages. What it does say, in plain guidance, is: don’t just rewrite what’s already out there — add something original.
What information gain means
Information gain is an information-theory and patent concept often used by SEOs; Google does not document a public ranking factor with a universal score under this name. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google patent: Contextual estimation of link information gain Google’s helpful-content guidance does ask whether content adds original information, reporting, research, or analysis. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Helpful content guidance
Picture someone researching a topic. They open the first result, read it, then open the second. If the second page repeats everything the first one said — just worded differently — they learned nothing new from it. If it gives them a fresh data point, a first-hand story, or a different angle, they gained information.
That’s the whole idea. Information gain is about novelty relative to what the searcher has already seen, not about how polished or complete a single page is on its own. A 4,000-word article that carefully restates the same consensus facts as the top ten results has low information gain. A short post with one genuinely new number can have high information gain.
Where the term comes from
Google was granted a patent (filed in 2018, granted in 2022) called “Contextual estimation of link information gain.” It describes scoring a document for how much additional information it contains beyond documents a user has already seen — and using that to help decide what to rank or show next.
Two things to keep straight:
- A patent is not confirmation. Google files thousands of patents. Having one proves Google could do something, not that it does. Google has never confirmed it uses this in live search ranking.
- It’s not a synonym for “quality.” The patent is narrowly about newness, not general goodness.
Why it matters more now
Search results are increasingly answered by AI — Google’s AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index., AI Mode, ChatGPT, and others. These systems are very good at summarizing what’s already been published everywhere. So if your page just says what everyone else already says, an AI can absorb it and answer the question without you.
The content that survives that is the content AI can’t just paraphrase: original research, your own data, first-hand experience, expert quotes no one else has. That’s information gain in practice.
What to do about it
- Before publishing, ask: what does this page say that the pages already ranking don’t? If the honest answer is “nothing,” you haven’t added information gain.
- Add something only you have — a small survey, a screenshot from a real test, a quote from someone on your team who actually does the work.
- Stop chasing “more” (more words, more subtopics, more FAQs). Chase new.
Want the patent details, the scope caveats, the official Google language that echoes this, and the practical playbook? Switch to the Advanced tab.
TL;DR — Information gainInformation gain is how much new information a page adds beyond what a searcher has already seen in prior results on the same topic — novelty relative to the existing corpus, not general content quality. The term comes from a granted Google patent that Google has never confirmed using in live ranking. is novelty relative to the corpus a searcher has already seen on a topic — not comprehensiveness, not E-E-A-T, not “quality” in the abstract. It’s named after a granted Google patent (“Contextual estimation of link information gain,” filed 2018, granted 2022) that scores documents 0.00–1.00 for how much new information they add and ranks partly on that. But: Google has never confirmed live use; the patent’s own framing is about what to show next after a user has seen some results, not the first page; and there’s no visible score, no API, no disclosed formula. What is official is Google’s helpful-content guidance asking for “original information, reporting, research, or analysis.” It matters more in an AI-answer SERP because LLMsA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4). synthesize consensus for free — only genuinely new information forces a citation.
What information gain actually is (and isn’t)
The cited patent describes possible methods, not proof that a specific production ranking system operates exactly that way. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google patent: Contextual estimation of link information gain Use original-value guidance as an editorial principle, not a measurable ranking guarantee. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Helpful content guidance
Most articles on this topic quietly conflate three different things. Keep them separate:
-
Information gain — the marginal new information a document adds relative to what the searcher has already seen on the topic. It’s a delta.
-
Comprehensiveness / topical depth — covering everything about a topic. A page can be maximally comprehensive and have zero information gain if every fact in it already exists on the pages ranking above it.
-
E-E-A-T / content quality — trust, expertise, experience, authority. Related, but a different judgment. You can be an unimpeachable expert and still publish a page that adds nothing new.
-
Truth and accuracy — being novel isn’t the same as being correct. A page can add a genuinely new claim and still get it wrong. Novelty, accuracy, and trust are separate judgments; treat originality as one input to publish, not a substitute for verifying it.
If you take one thing from this page: information gain is specifically about the delta, not the depth and not the trust. That distinction is the thing nearly every competing explainer blurs.
One more distinction worth naming, because SEO borrowed the term loosely: “information gain” also has an older, more technical meaning outside search. In information theory it traces back to Claude Shannon’s 1948 work on entropy and information measures — a mathematical account of how much a signal reduces uncertainty, unrelated to search rankings. In information retrieval research, a parallel idea shows up as novelty and diversity: methods like Maximal Marginal Relevance (Carbonell and Goldstein, 1998) rerank or select results to reduce redundancy against what a reader has already seen. Google’s patent draws on the same underlying intuition — new relative to what’s already been shown — but none of these are interchangeable. The SEO industry’s “information gain” is shorthand built on a specific patent, not a direct application of Shannon’s math or of published IR diversity algorithms, even though the instinct they share (don’t just repeat what’s already there) rhymes across all three.
The top box represents consensus facts A, B, and C that the reader has already seen. The low-gain candidate repeats A, B, and C in different wording. The high-gain candidate adds original evidence D that was absent from the earlier sources. The comparison is conceptual, not a confirmed ranking factor or public Google score.
© Patrick Stox LLC · CC BY 4.0 ·
The Google patent behind the term
What it says
The patent is “Contextual estimation of link information gain” (US20200349181A1, filed October 2018, granted June 2022; inventors include Victor Carbune and Pedro Gonnet Anders). The core definition, verbatim from the patent:
“An information gain score for a given document is indicative of additional information that is included in the given document beyond information contained in other documents that were already presented to the user.”
It describes a machine-learning model that takes semantic vectors from documents the user already viewed plus data from a candidate new document, and outputs, in the patent’s words:
“a quantitative score between 0.00 and 1.00, with 0.00 indicating that no information gain is to be expected”
— with 1.00 meaning the document contains only information not in what was already seen. Documents may then be ranked at least partly on their respective information gain scores.
What it doesn’t say
This is where most SEO content overreaches, so let me be precise about the limits:
- No confirmed live use. A granted patent is evidence Google could deploy this, not that it does. Google has not confirmed or denied using this mechanism in ranking. Anyone telling you it’s a confirmed ranking factor is guessing.
- Narrower scope than assumed. The patent’s framing centers on choosing what document or link to surface next — think a conversational or assistant follow-up after a user has already viewed some results — not on the first page of ten blue links. Roger Montti’s analysis at Search Engine Journal makes this scope point well: the emphasis is on automated assistants, and these scores aren’t described as applying to the first set of results (SEJ).
- No public formula. The patent names “semantic vectors” and salient extracted information, but discloses no weights, no feature list, and no reproducible scoring method. Do not trust any article that hands you a precise “how Google calculates it” formula — it’s invented.
- It’s a patent family, not one filing. Google was granted a continuation in the same family — US12013887B2, same title, same assignee (Google LLC), priority date back to the original October 2018 filing. A continuation means more disclosed claim language exists in the family, not that anything new about deployment has been confirmed — a granted patent and its continuation still only establish what Google disclosed it could build, not what’s running in production.
Bill Slawski’s early breakdown at Go Fish Digital framed the underlying problem Google is solving — that when many documents share a topic, they tend to contain similar information — and how the patent proposes ranking partly on information gain scores to address it (Go Fish Digital).
Google’s official guidance that echoes the idea
Here’s the honest part most explainers skip: no Google document, blog post, or spokesperson uses the phrase “information gain” in public search guidance. The term is industry shorthand. But Google’s actual official guidance — “Creating helpful, reliable, people-first content” — repeatedly describes the same underlying idea in its self-assessment questions. These are verbatim from that page:
“Does the content provide original information, reporting, research, or analysis?”
“If the content draws on other sources, does it avoid simply copying or rewriting those sources, and instead provide substantial additional value and originality?”
“Does the content provide insightful analysis or interesting information that is beyond the obvious?”
“Does the content provide substantial value when compared to other pages in search results?”
That last one — substantial value compared to other pages in search results — is about as close as Google’s official language gets to the information-gain concept without using the term. This is the strongest legitimately-official tie-in available, and it’s why “information gain” caught on as useful shorthand for a real principle Google clearly cares about, even though it’s not Google’s word.
Does Bing use information gain?
Not that I’ve foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore.. No Bing or Microsoft document or spokesperson uses the term in the material I’ve reviewed. Bing’s publicly discussed ranking factors — site and author reputation, content completeness, semantic relevance — point in a similar direction (depth and originality versus peers), but that’s an adjacent idea, not the same claim and not the same vocabulary. If someone quotes Bing using the literal phrase “information gain,” treat it with suspicion until you see the primary source.
Why information gain matters more in an AI-answer-heavy SERP
This is the part that makes information gain more than a patent-trivia curiosity.
AI has read the internet. ChatGPT, Claude, Gemini, and Google’s own AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. and AI Mode are excellent at one specific thing: synthesizing and repackaging what already exists across the web. Nathan Wahl at Animalz makes the argument sharply — these systems can already reiterate and repackage everything that’s been published, so the question becomes why publish anything that isn’t additive. When Google synthesizes an answer, it pulls from multiple sources — Wahl’s framing is that you don’t need to outrank giants if your content contains information theirs doesn’t (Animalz).
That’s the mechanism. If your page restates consensus, an AI Overview absorbs it and answers without sending anyone to you. If your page contains something only you have — a proprietary number, a real experiment, an expert’s first-hand take — an AI system has to draw on you specifically to include that fact, rather than paraphrasing it from wherever else it appears. That’s not a guarantee: retrieval, inclusion in the model’s context, generation, and citation are separate steps, and originality clearing one of them doesn’t mean it clears the rest. Scoped research on generative-engine visibility (Aggarwal et al., “GEO: Generative Engine Optimization”) finds that adding sources, statistics, and quotations to content can improve visibility in AILLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals.-generated answers in their tested setups — evidence that originality helps, not proof that it guarantees a citation, a ranking, or traffic. Information gain is a precondition for being cited instead of paraphrased, not a promise of it.
Bernard Huang at Clearscope has tied this to Google’s Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself.: aim for content that covers concepts and entities on the fringe of what Google already “knows” about a topic, because LLMs are great at consensus but genuine novelty is still where humans add value (Clearscope). Amanda King’s early Search Engine Land framing captured the core idea — an information gain score is essentially a measure of how unique your content is versus the rest of the corpus, and content risks being demoted if it lacks uniqueness even when it’s just the same ideas in different words (Search Engine Land). Andrew Holland later reframed the whole thing for the AI era: the term means different things to different people, but the job is to keep increasing the rate of information gain — and increasingly we’re optimizing for AI, not just for Google (Search Engine Land).
I’ve seen this from the other direction, too. I ran some tests trying to rank content in Google’s AI Mode, and the content simply didn’t rank — even though it was, as far as I could tell, better and more relevant than pages that did. One of the possibilities I floated publicly was a lack of information gain in the articles: if the content is a well-written restatement of what’s already out there, there may be nothing for the system to reward (my post on X).
What actually produces information gain
Synthesizing across the sources — and my own experience — here’s what genuinely moves the needle:
- Original research, surveys, and experiments. The most defensible form. Run a study, publish the data, let others cite you. At Ahrefs, our best-performing content skews heavily toward data studies for exactly this reason.
- Proprietary / first-party data. Internal product or usage data nobody else can reproduce. If you have a dataset, that dataset is your information gain.
- Expert interviews and first-hand experience. This is my go-to tactic. Rather than shipping unsourced generated text, message a handful of real subject-matter experts you already have access to — over Slack, email, or a quick AI-assisted call — a few times a week, and capture their actual stories and experience. That’s original information that didn’t exist on the web before you published it. I applied this to my own rebuild of Ahrefs’ technical SEOTechnical SEO is the practice of making a site easy for search engines to crawl, render, index, and (now) be eligible for AI answers. It's the foundation that lets your content and links rank — not a ranking trick of its own. hub (roughly 160 AI-generated pages) by layering expert review and reader feedback on top so the pages carry something beyond regurgitated consensus.
- Contrarian or updated takes on consensus. Testing a widely repeated claim and reporting what actually happened is information gain, even when the “study” is small.
- Building on a predecessor’s work. Take someone’s published research, extend it, add the next data point. You’re adding to the corpus, not copying it.
Ahrefs’ own content-quality ladder (Si Quan Ong’s “How to Create Quality Content”) is a useful map here: it climbs from simple listicles up to research studies and original ideas — effectively an information-gain ladder, prizing first-hand data over aggregation, without ever using the term.
How to evaluate your own content for information gain
Adapt Google’s helpful-content questions into a pre-publish gut check:
- Open the top results for your target query. List what each one already says.
- For your draft, highlight every sentence that adds a fact, number, angle, or experience not on that list. If nothing is highlighted, you have a rewrite, not a resource.
- Ask whether an AI could answer the query fully from the pages already ranking. If yes, your only path to relevance is contributing something those pages lack.
- Sanity-check the “value compared to other pages in search results” question literally — not “is this good,” but “is this additive.”
The bottom line
Information gain is a real, useful concept sitting on top of an unconfirmed mechanism. Don’t sell it internally as a confirmed Google ranking factor with a score you can dial — that’s not true, and it’ll burn your credibility. Sell it as the thing that’s demonstrably working in an AI-heavy SERP: stop publishing better-optimized versions of what already exists, and start publishing what only you can. That’s the moat.
Related reading in this cluster: entity SEOEntity SEO is the practice of helping search engines and AI systems clearly identify, classify, and trust the entities you represent — your brand, your people, your products — rather than just matching keyword strings. The goal is to be an unambiguous, well-corroborated entity in machine knowledge systems so AI can cite you with confidence., schema markup for AISchema markup (structured data) is machine-readable code — usually JSON-LD — that labels what your content means using the schema.org vocabulary. For AI search it's infrastructure for entity disambiguation, not a direct citation lever: controlled studies found no meaningful uplift in AI citations from adding it., GEOGenerative Engine Optimization (GEO) is the practice of optimizing content and brand presence so AI-powered search engines and assistants — Google AI Overviews, ChatGPT, Perplexity — cite, recommend, or mention you when generating answers. Google's position is that it's still SEO., and AEOAnswer Engine Optimization (AEO) is the practice of structuring content so engines deliver it as a direct answer — featured snippets, voice assistants, and AI search — rather than just a ranked link. Coined for voice search in 2018 and revived for the LLM era. Google's position is that it's still SEO. all connect to how you get surfaced and cited once your content actually has something new to say.
AI summary
A condensed take on the Advanced version:
- Information gainInformation gain is how much new information a page adds beyond what a searcher has already seen in prior results on the same topic — novelty relative to the existing corpus, not general content quality. The term comes from a granted Google patent that Google has never confirmed using in live ranking. = novelty relative to what a searcher has already seen on a topic. It’s the marginal delta, not comprehensiveness and not E-E-A-T. A thorough page can have zero information gain.
- The name comes from a granted Google patent (“Contextual estimation of link information gain,” filed 2018, granted 2022) that scores documents 0.00–1.00 for how much new information they add and ranks partly on that.
- Google has never confirmed live use. The patent’s own framing is about what to show next after a user has seen some results — not the first page. There’s no visible score, no API, and no disclosed formula; any “information gain score” in a tool is an approximation.
- What’s official is Google’s helpful-content guidance asking whether content provides “original information, reporting, research, or analysis” and “substantial value when compared to other pages in search results” — the same idea, without the term.
- Why it matters now: AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index. and AI Mode synthesize consensus content for free. Genuinely new information — original research, proprietary data, first-hand expertise — is what an AI system has to draw on you for instead of paraphrasing elsewhere, though that’s not a guarantee of citation, ranking, or traffic (retrieval, inclusion, generation, and citation are separate steps).
- Novelty isn’t accuracy. A page can add a genuinely new claim and still be wrong — treat originality as one input to publish, not a substitute for verifying it.
- What produces it: original studies, first-party data, expert interviews and first-hand experience, contrarian/updated takes, extending others’ research.
- Beware fabricated specifics — precise “core update confirmed information gain, +X% visibility” stats with no primary source are a myth pattern, not data.
Official documentation
Primary sources. Note that Google’s guidance never uses the term “information gain” — the patent does, but a patent is a legal filing, not a statement about live ranking.
Google — the patent (primary source, not “guidance”)
- Contextual estimation of link information gain (US20200349181A1) — the filing the term is named after: the 0.00–1.00 scoring description and the “additional information beyond documents already presented to the user” definition.
Google — official helpful-content guidance (uses the concept, not the term)
- Creating helpful, reliable, people-first content — the self-assessment questions about original information/research/analysis and substantial value versus other results. This is the closest official language to information gainInformation gain is how much new information a page adds beyond what a searcher has already seen in prior results on the same topic — novelty relative to the existing corpus, not general content quality. The term comes from a granted Google patent that Google has never confirmed using in live ranking..
Quotes from the source
On-the-record language. The Google items below are verbatim from primary sources (the patent and Google’s helpful-content docs). Industry voices are paraphrased in the article body rather than quoted, because I couldn’t independently re-verify their exact wording against the original pages — I’m not going to put words in quotation marks I can’t confirm.
Google — the patent (US20200349181A1)
- “An information gainInformation gain is how much new information a page adds beyond what a searcher has already seen in prior results on the same topic — novelty relative to the existing corpus, not general content quality. The term comes from a granted Google patent that Google has never confirmed using in live ranking. score for a given document is indicative of additional information that is included in the given document beyond information contained in other documents that were already presented to the user.” — Contextual estimation of link information gain. Read the patent
- “a quantitative score between 0.00 and 1.00, with 0.00 indicating that no information gain is to be expected” — the patent’s described scoring scale (1.00 meaning the document contains only information not already seen). Read the patent
Google — helpful-content guidance (the concept, without the term)
- “Does the content provide original information, reporting, research, or analysis?” Jump to quote
- “If the content draws on other sources, does it avoid simply copying or rewriting those sources, and instead provide substantial additional value and originality?” Jump to quote
- “Does the content provide insightful analysis or interesting information that is beyond the obvious?” Jump to quote
- “Does the content provide substantial value when compared to other pages in search results?” Jump to quote
Playbook: adding real information gain
A repeatable process for turning a “yet another article on X” draft into something additive.
1. Map the corpus first. Open the top 10 results for your query. In a doc, list the distinct claims, data points, and angles each one already covers. This is the baseline the searcher will have “already seen.”
2. Find the gap. Where is the consensus thin, outdated, or untested? What question do the ranking pages raise but not answer? That gap is your target.
3. Pick your source of novelty. In rough order of durability:
- Original research / a survey / an experiment you run.
- First-party or proprietary data you already have.
- Expert interviews and first-hand experience.
- A contrarian or updated take you can actually support.
4. The expert-interview loop (my default). You almost always have subject-matter experts within reach. A few times a week, message a handful of them — Slack, email, or a short AI-assisted call — with one sharp question. Capture the story, the number, the “actually, in practice…” detail. Attribute it. That’s original information that didn’t exist on the web before you published it.
5. Attribute everything. Source your claims. Link out. Show your data. Content that sources where things came from reads as additive; unsourced generated text reads as filler — and increasingly gets treated that way.
6. Pre-publish check. Highlight every sentence in your draft that isn’t already on one of the top-10 pages. If little is highlighted, don’t ship it yet — go back to step 3.
7. Layer, don’t dump. If you’re generating content at scale, don’t stop at the generation. Layer expert review, reader feedback, and real examples on top so each page carries something beyond consensus. Volume without novelty is the thing AI absorbs for free.
Anti-patterns: how “information gain” goes wrong
Treating it as a confirmed ranking factor with a score you can dial. It’s a granted patent Google has not confirmed using in live ranking. There’s no Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. field, no log signal, no public API. Any “information gainInformation gain is how much new information a page adds beyond what a searcher has already seen in prior results on the same topic — novelty relative to the existing corpus, not general content quality. The term comes from a granted Google patent that Google has never confirmed using in live ranking. score” in a tool is that tool’s approximation, not Google’s. Don’t pitch it internally as a confirmed lever — you’ll lose credibility when someone asks for the source.
Repeating suspiciously precise, uncited stats. A wave of 2026-dated posts assert things like “the [date] core update confirmed information gain as the dominant signal, +15–25% visibility for original content, –30–50% for templated content,” down to specific timestamps. No primary source substantiates any of it. These read as fabricated or hallucinated specifics riding a buzzword. If a stat about information gain has no traceable primary source, don’t repeat it.
Confusing “write more” with information gain. Length and comprehensiveness are not information gain. A 5,000-word article that rehashes consensus has less of it than a 400-word post with one new data point. The FAQ-stuffing move — bolt 50 or 100 FAQs onto a page — is the clearest example: it adds words, not novelty, and it doesn’t work.
Conflating it with E-E-A-T or “quality.” Related but distinct. You can be a genuine expert and still publish a page that adds nothing new to the corpus. Information gain is the marginal delta, not the trust/expertise judgment.
Assuming it applies to the first page of results. The patent’s own framing centers on what to surface next — assistant/follow-up contexts — after a user has seen some results, not necessarily the initial ten blue links. Most explainers skip this scope caveat; don’t.
Claiming Google or Bing said “information gain.” Neither has used the literal phrase in public ranking guidance that I’ve foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore.. It’s industry shorthand built on a patent (Google) and general helpful-content principles. Treat any attributed “information gain” quote from Google or Bing with suspicion until you see the primary source.
Prompts for information-gain reviews
Compare my draft with these competing-page notes. Build a claim matrix with one row
per substantive point and columns for: already common in the corpus, genuinely new,
new only in wording, supported by first-hand evidence, source needed, and remove or
keep. Do not assign a fake information-gain score. Finish with the three strongest
original contributions and the evidence needed to publish each.
DRAFT:
[paste]
COMPETING-PAGE NOTES:
[paste]Turn these raw first-party materials into an evidence plan, not finished prose.
Separate proprietary data, direct observation, expert experience, original examples,
and repeatable analysis. For each candidate contribution, state the claim it could
support, the minimum method disclosure, limitations, and what would make the claim too
weak to publish. Do not infer results that are absent.
[paste interview notes, dataset description, tests, or case documentation] Frameworks for creating real information gain
The corpus–claim–evidence framework
- Corpus: Summarize what the current useful pages already agree on.
- Claim: State what your page would add beyond that consensus.
- Evidence: Name the data, observation, experience, or source that makes the new claim defensible.
- Boundary: Document where the evidence does not generalize.
If the claim changes only phrasing, it is not new. If the evidence cannot support the claim, it is not ready.
The novelty ladder
- Restatement: Same facts, different wording.
- Synthesis: Existing facts connected in a useful new way.
- Application: Existing knowledge tested in a specific context.
- Observation: First-hand evidence others do not have.
- Discovery: A reproducible finding that changes what the reader knows.
The ladder is a planning model, not Google’s hidden score.
Information gain cheat sheet
| Candidate addition | Adds useful novelty? | Proof needed |
|---|---|---|
| Longer explanation of the same facts | Usually no | None; tighten it |
| Clear synthesis across credible sources | Sometimes | Traceable source map |
| Original expert interview | Yes, if substantive | Speaker, role, accurate notes or recording |
| Proprietary dataset finding | Yes, if reproducible | Method, sample, definitions, limitations |
| First-hand implementation lesson | Yes, when specific | Context, observed result, boundaries |
| AI-generated examples presented as real | No | Do not publish as evidence |
| Patent language described as a live ranking factor | No | Avoid the unsupported leap |
Reality check: Google exposes no public information-gain score or formula. Evaluate the contribution and its evidence, not a made-up decimal.
Resources worth your time
My writing & speaking
- Does AI Search Traffic Convert Better Than Traditional Search? (Ahrefs) — the data behind why AI-search visibility is worth chasing, which is the reason information gainInformation gain is how much new information a page adds beyond what a searcher has already seen in prior results on the same topic — novelty relative to the existing corpus, not general content quality. The term comes from a granted Google patent that Google has never confirmed using in live ranking. matters commercially.
- “GEO? AEO? LLMO? What’s With All This AI SEO Stuff?” — my Ahrefs Evolve 2025 deck on AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity., where the “publish what only you have” argument lives.
- My AI Mode ranking-test post (X) — the natural experiment where content that didn’t rank in AI Mode may have lacked information gain.
- Patrick Stox on Building in the GEO Era (Unscripted SEO Podcast) — where I lay out the “message five real experts, source your claims” tactic.
- Cutting-Edge AEO Strategies with Patrick Stox (Marketing Speak) — the episode framing information gain as the real moat, with the Ahrefs data-studies point.
- How to Create Quality Content (Ahrefs, by Si Quan Ong) — house content: the quality ladder from listicles up to original research, effectively an information-gain ladder.
From around the industry
- Contextual estimation of link information gain (US20200349181A1) — the actual patent. Read the source before you trust anyone’s summary of it.
- Creating helpful, reliable, people-first content (Google Search Central) — the official self-assessment questions that describe the concept without the term.
- Google’s Information Gain Patent For Ranking Web Pages (Roger Montti, Search Engine Journal) — the most careful piece on scope: assistant/follow-up context, not the first page.
- What is information gain in SEO & why it matters (Amanda King, Search Engine Land) — the early mainstream definition.
- Information gain: what this SEO “buzzword” really means (Andrew Holland, Search Engine Land) — the AI-era reframe.
- Information Gain in SEO (Bernard Huang, Clearscope) — the Knowledge GraphThe Knowledge Graph is Google's database of entities — people, places, organizations, and things — and the factual relationships between them. It's separate from any single website's structured data: your schema markup is one of many possible inputs to the graph, not the graph itself. framing and practical viewpoints.
- Information Gain: The SEO Theory that AI Made Mandatory (Nathan Wahl, Animalz) — the strongest “why now” argument tied to AI synthesis and multi-source citation.
- Information Gain Scores (Bill Slawski, via Go Fish Digital) — the original technical breakdown of the patent.
- What Is Information Gain in SEO & Does Google Measure It? (Rachel Handley, Semrush) — a solid additional definitional overview.
Test yourself: Information Gain
Five quick questions on what information gainInformation gain is how much new information a page adds beyond what a searcher has already seen in prior results on the same topic — novelty relative to the existing corpus, not general content quality. The term comes from a granted Google patent that Google has never confirmed using in live ranking. is, where it comes from, and what to do about it. Pick an answer for each, then check.
Information Gain
Information gain is how much new information a page adds beyond what a searcher has already seen in prior results on the same topic — novelty relative to the existing corpus, not general content quality. The term comes from a granted Google patent that Google has never confirmed using in live ranking.
Related: Entity SEO, Generative Engine Optimization (GEO), Answer Engine Optimization (AEO), Helpful Content Update
Information Gain
Information gain measures how much new information a document contributes beyond what a searcher has already encountered in earlier results for the same topic. It is not a synonym for “good content” or “comprehensive content” — a long, well-optimized page can have zero information gain if every fact in it already exists on the pages already ranking. Information gain is specifically about the marginal delta: does your page say something the top results do not?
The phrase comes from a real Google patent, “Contextual estimation of link information gain” (filed October 2018, granted June 2022). The patent describes scoring a document from 0.00 (nothing new beyond what the user already saw) to 1.00 (entirely new information) and ranking or selecting documents partly on that score. Importantly, Google has never confirmed this mechanism is live in ranking, and its framing centers on choosing what to show next after a user has already viewed some results — not the first page of blue links.
There is no visible “information gain score” in Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results., no log field, and no public API. Any score shown in a third-party tool is that tool’s own approximation. The idea is still directionally supported by Google’s public helpful-content guidance, which repeatedly asks whether content provides original information, reporting, research, or analysis — the same underlying principle without ever using the patent’s terminology.
Information gain matters more in an AI-answer-heavy SERP: large language modelsA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4). can already synthesize and repackage consensus information, so content that merely restates what’s known adds little that a machine can’t summarize. Original research, proprietary data, and first-hand expertise are what an AI system has to draw on specifically rather than paraphrase from elsewhere — though that’s not a guarantee of citation, ranking, or traffic; retrieval, inclusion, generation, and citation are separate steps.
Related: Entity SEO, Generative Engine Optimization (GEO), Answer Engine Optimization (AEO), Helpful Content Update
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Updated Jul 18, 2026.
Editorial summary and recorded change details.Summary
Added the patent-family continuation, distinguished SEO's 'information gain' from Shannon information theory and IR novelty/diversity research, and hedged the AI-citation claim so it no longer implies a guarantee.
Change details
-
Noted the granted continuation patent US12013887B2 (same family, same title and assignee, priority back to the October 2018 filing) alongside the original filing, with the same not-evidence-of-deployment caveat.
-
Added a distinction between the SEO industry's 'information gain' shorthand and the separate, older concepts of Shannon information theory and IR novelty/diversity research (e.g., Maximal Marginal Relevance).
-
Added a novelty-is-not-accuracy distinction: a genuinely new claim can still be wrong.
-
Reworded the AI-citation claim so it no longer states an AI 'has to' cite original content; cited scoped GEO (Generative Engine Optimization) research showing originality can help visibility without guaranteeing citation, ranking, or traffic, and applied the same hedge to the glossary entry and AI summary.
Full comparison unavailable — no prior snapshot was archived for this revision.