Guide : Information Gain

Ce que information gain signifie in SEO and AI search — how Google's Information Gain Score patent fonctionne, pourquoi original research and unique données outperform regurgitated content, and how to mesurer it.

Première publication : 2 juil. 2026 · Dernière mise à jour : 3 août 2026 · Advanced
Langues

Information gain is how beaucoup *nouveau* information votre page adds au-delà ce que a searcher has déjà seen in le résultats pour a topic — novelty relative to the existing corpus, pas general content quality. The term comes from a réel Google patent ('Contextual estimation of lien information gain,' filed 2018, granted 2022) que scores documents 0,00–1,00 pour novelty and ranks partly on que score — but Google has jamais confirmed en utilisant it in live ranking, and the patent's framing is à propos de ce que to montrer *suivant* après a utilisateur has seen some results, pas the premier page. There's aucun visible score in Search Console, aucun API, and aucun public formula: quelconque 'information gain score' in a outil is an approximation. What's réel and official is Google's helpful-content guidance asking si content offers 'original information, reporting, research, or analysis.' It matters plus now parce que AI Overviews and AI Mode peut déjà synthesize consensus content — seulement genuinely nouveau information (proprietary données, original research, first-hand expertise) is ce que AI has to cite au lieu de absorb.

TL;DR — Information gain is novelty relative to the corpus a searcher has déjà seen on a topic — pas comprehensiveness, pas E-E-A-T, pas “quality” in the abstract. It’s named après a granted Google patent (“Contextual estimation of lien information gain,” filed 2018, granted 2022) que scores documents 0,00–1,00 pour how beaucoup nouveau information ils ajouter and ranks partly on que. But: Google has jamais confirmed live utiliser; the patent’s propre framing is à propos de ce que to montrer suivant après a utilisateur has seen some results, pas the premier page; and there’s aucun visible score, aucun API, aucun disclosed formula. Ce que is official is Google’s helpful-content guidance asking pour “original information, reporting, research, or analysis.” It matters plus in an AI-answer SERP parce que LLMs synthesize consensus pour free — seulement genuinely nouveau information forces a citation.

Ce que information gain en réalité is (and isn’t)

The cited patent describes possible méthodes, pas proof que a spécifique production ranking system operates exactly que façon. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google patent: Contextual estimation of link information gain Utiliser original-value guidance as an editorial principle, pas a measurable ranking guarantee. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Helpful content guidance

La plupart articles on ce topic quietly conflate three différent choses. Garder les separate:

  • Information gain — the marginal nouveau information a document adds relative to ce que the searcher has déjà seen on the topic. It’s a delta.

  • Comprehensiveness / topical depth — covering everything à propos de a topic. Une page peut be maximally comprehensive and have zero information gain si every fact in it déjà exists on lune pages ranking ci-dessus it.

  • E-E-A-T / content quality — trust, expertise, experience, authority. Connexe, but a différent judgment. Vous pouvez be an unimpeachable expert and encore publish a page que adds nothing nouveau.

  • Truth and accuracy — being novel isn’t the même as being correct. Une page peut ajouter a genuinely nouveau claim and encore obtenir it incorrect. Novelty, accuracy, and trust are separate judgments; treat originality as un input to publish, pas a substitute pour verifying it.

Si vous prendre un chose from ce page: information gain is specifically à propos de the delta, pas the depth and pas the trust. Que distinction is the chose nearly every competing explainer blurs.

Un plus distinction worth naming, parce que SEO borrowed the term loosely: “information gain” aussi has an older, plus technical meaning outside search. In information theory it traces back to Claude Shannon’s 1948 fonctionner on entropy and information measures — a mathematical account of how beaucoup a signal reduces uncertainty, unrelated to search rankings. In information retrieval research, a parallel idea montre up as novelty and diversity: méthodes comme Maximal Marginal Relevance (Carbonell and Goldstein, 1998) rerank or select results to reduce redundancy contre ce que a reader has déjà seen. Google’s patent draws on the même underlying intuition — nouveau relative to what’s déjà been affiché — but none of ces are interchangeable. The SEO industry’s “information gain” is shorthand construit on a spécifique patent, pas a direct application of Shannon’s math or of publié IR diversity algorithms, même though the instinct ils share (don’t simplement repeat what’s déjà là) rhymes à travers tout three.

Information gain is the delta beyond what was already available—not the total length or polish of the candidate page. Source : /ai-search/optimization/information-gain/

The top box represents consensus facts A, B, and C that the reader has already seen. The low-gain candidate repeats A, B, and C in different wording. The high-gain candidate adds original evidence D that was absent from the earlier sources. The comparison is conceptual, not a confirmed ranking factor or public Google score.

© Patrick Stox LLC · CC BY 4.0 ·

The Google patent behind the term

Ce que it dit

The patent is “Contextual estimation of link information gain” (US20200349181A1, filed October 2018, granted June 2022; inventors inclure Victor Carbune and Pedro Gonnet Anders). The core definition, verbatim from the patent:

“An information gain score pour a donné document is indicative of additional information que is inclus in the donné document au-delà information contained in autre documents que were déjà presented to the utilisateur.”

It describes a machine-learning model que takes semantic vectors from documents the utilisateur déjà viewed plus données from a candidate nouveau document, and outputs, in the patent’s words:

“a quantitative score entre 0,00 and 1,00, with 0,00 indicating que aucun information gain is to be attendu”

— with 1,00 meaning the document contient seulement information pas in ce que was déjà seen. Documents may alors be ranked au moins partly on leur respective information gain scores.

Ce que it doesn’t dire

Ce is où la plupart SEO content overreaches, so let me be precise à propos de the limites:

  • Aucun confirmed live utiliser. A granted patent is evidence Google pourrait deploy ce, pas que it fait. Google has pas confirmed or denied en utilisant ce mechanism in ranking. Anyone telling vous it’s a confirmed ranking factor is guessing.
  • Narrower scope que assumed. The patent’s framing centers on choosing ce que document or lien to surface suivant — think a conversational or assistant follow-up après a utilisateur has déjà viewed some results — pas on the premier page of ten blue liens. Roger Montti’s analysis at Moteur de recherche Journal rend ce scope point bien: the emphasis is on automated assistants, and ces scores aren’t décrit as applying to the premier définir of results (SEJ).
  • Aucun public formula. The patent noms “semantic vectors” and salient extracted information, but discloses aucun weights, aucun fonctionnalité liste, and aucun reproducible scoring méthode. Ne faites pas trust quelconque article que hands vous a precise “how Google calculates it” formula — it’s invented.
  • It’s a patent family, pas un filing. Google was granted a continuation in the même family — US12013887B2, même title, même assignee (Google LLC), priority date back to the original October 2018 filing. A continuation signifie plus disclosed claim language exists in the family, pas que anything nouveau à propos de deployment has been confirmed — a granted patent and its continuation encore seulement establish ce que Google disclosed it pourrait construire, pas what’s running in production.

Bill Slawski’s early breakdown at Go Fish Digital framed the underlying problem Google is solving — que quand nombreux documents share a topic, ils tend to contain similaire information — and how the patent proposes ranking partly on information gain scores to adresse it (Go Fish Digital).

Google’s official guidance que echoes the idea

Here’s the honest partie la plupart explainers skip: aucun Google document, blog post, or spokesperson uses the phrase “information gain” in public search guidance. The term is industry shorthand. But Google’s réel official guidance — “Creating helpful, reliable, people-first content” — repeatedly describes the même underlying idea in its self-assessment questions. Ces are verbatim from que page:

“Fait le contenu provide original information, reporting, research, or analysis?”

“Si le contenu draws on autre sources, fait it éviter simply copying or rewriting ceux sources, and à la place provide substantial additional valeur and originality?”

“Fait le contenu provide insightful analysis or interesting information que is au-delà the obvious?”

“Fait le contenu provide substantial valeur quand comparé to autre pages in résultats de recherche?”

Que dernier un — substantial valeur comparé to autre pages in résultats de recherche — is à propos de as fermer as Google’s official language obtient to the information-gain concept sans en utilisant the term. Ce is the strongest legitimately-official tie-in disponible, and it’s pourquoi “information gain” caught on as utile shorthand pour a réel principle Google clearly cares à propos de, même though it’s pas Google’s word.

Fait Bing utiliser information gain?

Pas que I’ve trouvé. Aucun Bing or Microsoft document or spokesperson uses the term in the material I’ve reviewed. Bing’s publicly discussed ranking factors — site and author reputation, content completeness, semantic relevance — point in a similaire direction (depth and originality versus peers), but that’s an adjacent idea, pas the même claim and pas the même vocabulary. Si someone quotes Bing en utilisant the literal phrase “information gain,” treat it with suspicion jusqu’à vous voir the principal source.

Pourquoi information gain matters plus in an AI-answer-heavy SERP

Ce is the partie que rend information gain plus que a patent-trivia curiosity.

AI has lire the internet. ChatGPT, Claude, Gemini, and Google’s propre AI Overviews and AI Mode are excellent at un spécifique chose: synthesizing and repackaging ce que déjà exists à travers the web. Nathan Wahl at Animalz rend the argument sharply — ces systems peut déjà reiterate and repackage everything that’s been publié, so the question becomes pourquoi publish anything que isn’t additive. Quand Google synthesizes an réponse, it pulls from multiple sources — Wahl’s framing is que vous don’t besoin to outrank giants si votre content contient information theirs doesn’t (Animalz).

That’s the mechanism. Si votre page restates consensus, an AI Overview absorbs it and réponses sans sending anyone to vous. Si votre page contient something seulement vous have — a proprietary number, a réel experiment, an expert’s first-hand prendre — an AI system has to draw on vous specifically to inclure que fact, plutôt que paraphrasing it from wherever sinon it apparaît. That’s pas a guarantee: retrieval, inclusion in the model’s context, generation, and citation are separate steps, and originality clearing un of les doesn’t mean it clears the rest. Scoped research on generative-engine visibility (Aggarwal et al., “GEO: Generative Engine Optimization”) trouve que ajout sources, statistics, and quotations to content peut améliorer visibility in AI-generated réponses in leur testé setups — evidence que originality helps, pas proof que it guarantees a citation, a ranking, or trafic. Information gain is a precondition pour being cited au lieu de paraphrased, pas a promise of it.

Bernard Huang at Clearscope has tied ce to Google’s Knowledge Graph: aim pour content que covers concepts and entities on the fringe of ce que Google déjà “knows” à propos de a topic, parce que LLMs are great at consensus but genuine novelty is encore où humans ajouter valeur (Clearscope). Amanda King’s early Moteur de recherche Land framing captured the core idea — an information gain score is essentially a mesurer of how unique votre content is versus the rest of the corpus, and content risks being demoted si it lacks uniqueness même quand it’s simplement the même ideas in différent words (Moteur de recherche Land). Andrew Holland plus tard reframed the whole chose pour the AI era: the term signifie différent choses to différent personnes, but the job is to garder increasing the rate of information gain — and increasingly we’re optimizing pour AI, pas simplement pour Google (Moteur de recherche Land).

I’ve seen ce from the autre direction, aussi. I ran some tests trying to rank content in Google’s AI Mode, and le contenu simply didn’t rank — même though it was, as far as I pourrait tell, meilleur and plus relevant que pages que did. Un of the possibilities I floated publicly was a lack of information gain in the articles: si le contenu is a well-written restatement of what’s déjà out là, là may be nothing pour the system to reward (my post on X).

Ce que en réalité produces information gain

Synthesizing à travers the sources — and my propre experience — here’s ce que genuinely moves the needle:

  • Original research, surveys, and experiments. The la plupart defensible formulaire. Run a study, publish the données, let others cite vous. At Ahrefs, our best-performing content skews heavily toward données studies pour exactly ce raison.
  • Proprietary / first-party données. Internal product or usage données nobody sinon peut reproduce. Si vous have a dataset, que dataset is votre information gain.
  • Expert interviews and first-hand experience. Ce is my go-to tactic. Plutôt que shipping unsourced generated text, message a handful of réel subject-matter experts vous déjà have accès to — over Slack, email, or a rapide AI-assisted appel — a few times a week, and capture leur réel stories and experience. That’s original information que didn’t exist on the web avant vous publié it. I applied ce to my propre rebuild of Ahrefs’ SEO technique hub (roughly 160 AI-generated pages) by layering expert examiner and reader feedback on top so the pages carry something au-delà regurgitated consensus.
  • Contrarian or mis à jour takes on consensus. Testing a widely repeated claim and reporting ce que en réalité happened is information gain, même quand the “study” is petit.
  • Building on a predecessor’s fonctionner. Prendre someone’s publié research, extend it, ajouter the suivant données point. You’re ajout to the corpus, pas copying it.

Ahrefs’ propre content-quality ladder (Si Quan Ong’s “How to Create Quality Content”) is a utile map ici: it climbs from simple listicles up to research studies and original ideas — effectively an information-gain ladder, prizing first-hand données over aggregation, sans ever en utilisant the term.

How to evaluate votre propre content pour information gain

Adapt Google’s helpful-content questions into a pre-publish gut vérifier:

  • Ouvrir the top results pour votre target requête. Liste ce que chaque un déjà dit.
  • Pour votre draft, highlight every sentence que adds a fact, number, angle, or experience pas on que liste. Si nothing is highlighted, vous have a rewrite, pas a resource.
  • Demander si an AI pourrait réponse the requête entièrement from lune pages déjà ranking. Si yes, votre seulement chemin to relevance is contributing something ceux pages lack.
  • Sanity-check the “value compared to other pages in search results” question literally — pas “is this good,” but “is this additive.”

L’essentiel

Information gain is a réel, utile concept sitting on top of an unconfirmed mechanism. Don’t sell it internally as a confirmed Google ranking factor with a score vous pouvez dial — that’s pas vrai, and it’ll burn votre credibility. Sell it as the chose that’s demonstrably working in an AI-heavy SERP: arrêter publishing better-optimized versions of ce que déjà exists, and commencer publishing ce que seulement vous pouvez. That’s the moat.

Connexe reading in ce cluster: entity SEO, balisage de données structurées pour AI, GEO, and AEO tout connecter to how vous obtenir surfaced and cited une fois votre content en réalité has something nouveau to dire.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.