Poradnik: Information Gain

co information gain means in SEO i wyszukiwanie AI — how Google's Information Gain Score patent działa, why original research i unique data outperform regurgitated treść, i how to mierzyć it.

Opublikowano po raz pierwszy: 2 lip 2026 · Ostatnia aktualizacja: 3 sie 2026 · Advanced
Języki

Information gain jest how much *new* information twój strona dodaje beyond co a searcher ma już seen in the wyniki dla a topic — novelty relative to the existing corpus, nie general jakość treści. The term comes z a rzeczywisty Google patent ('Contextual estimation of link information gain,' filed 2018, granted 2022) że scores documents 0,00–1,00 dla novelty i ranks partly on że score — ale Google ma nigdy confirmed używając it in live ranking, i the patent's framing jest o co to pokazywać *następny* po a użytkownik ma seen niektóre wyniki, nie the pierwszy strona. There's no visible score in Search Console, no API, i no public formula: dowolny 'information gain score' in a narzędzie jest an approximation. co's rzeczywisty i official jest Google's pomocny-treść guidance asking whether treść oferty 'original information, reporting, research, lub analiza.' It matters więcej now ponieważ Omówienia AI i AI Mode może już synthesize consensus treść — tylko genuinely new information (proprietary data, original research, pierwszy-hand expertise) jest co AI ma to cite zamiast absorb.

TL;DR — Information gain jest novelty relative to the corpus a searcher ma już seen on a topic — nie comprehensiveness, nie E-E-A-T, nie “quality” in the abstract. It’s named po a granted Google patent (“Contextual estimation of link information gain,” filed 2018, granted 2022) że scores documents 0,00–1,00 dla how much new information they dodawać i ranks partly on że. ale: Google ma nigdy confirmed live używać; the patent’s own framing jest o co to pokazywać następny po a użytkownik ma seen niektóre wyniki, nie the pierwszy strona; i there’s no visible score, no API, no disclosed formula. co jest official jest Google’s pomocny-treść guidance asking dla “original information, reporting, research, or analysis.” It matters więcej in an AI-answer SERP ponieważ LLMs synthesize consensus dla free — tylko genuinely new information forces a cytowanie.

co information gain actually jest (i isn’t)

The cited patent describes possible metody, nie proof że a specific production ranking system operates exactly że way. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google patent: Contextual estimation of link information gain używać original-wartość guidance as an editorial principle, nie a measurable ranking guarantee. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Helpful content guidance

najbardziej artykuły on ten topic quietly conflate three różny things. zachować them oddzielny:

  • Information gain — the marginal new information a document dodaje relative to co the searcher ma już seen on the topic. It’s a delta.

  • Comprehensiveness / topical depth — covering everything o a topic. A strona może być maximally comprehensive i mieć zero information gain if każdy fakt in it już exists on the strony ranking above it.

  • E-E-A-T / jakość treści — trust, expertise, experience, organ. powiązany, ale a różny judgment. You może być an unimpeachable expert i nadal publikować a strona że dodaje nothing new.

  • Truth i dokładność — będąc novel isn’t the same as będąc poprawny. A strona może dodawać a genuinely new twierdzenie i nadal get it błędny. Novelty, dokładność, i trust są oddzielny judgments; treat originality as one input to publikować, nie a substitute dla verifying it.

If you take one thing z ten strona: information gain jest specifically o the delta, nie the depth i nie the trust. że distinction jest the thing nearly każdy competing explainer blurs.

One więcej distinction worth naming, ponieważ SEO borrowed the term loosely: “information gain” również ma an older, więcej technical meaning outside search. In information theory it traces back to Claude Shannon’s 1948 działać on entropy i information measures — a mathematical account of how much a signal reduces uncertainty, unrelated to search rankings. In information pobieranie research, a parallel idea pokazuje up as novelty i diversity: metody like Maximal Marginal Relevance (Carbonell i Goldstein, 1998) rerank lub select wyniki to reduce redundancy wobec co a czytelnik ma już seen. Google’s patent draws on the same underlying intuition — new relative to co’s już był shown — ale none of te są interchangeable. The SEO industry’s “information gain” jest shorthand built on a specific patent, nie a bezpośredni application of Shannon’s math lub of opublikowany IR diversity algorithms, even though the instinct they share (don’t just repeat co’s już there) rhymes w całym wszystkie three.

Information gain is the delta beyond what was already available—not the total length or polish of the candidate page. Źródło: /ai-search/optimization/information-gain/

The top box represents consensus facts A, B, and C that the reader has already seen. The low-gain candidate repeats A, B, and C in different wording. The high-gain candidate adds original evidence D that was absent from the earlier sources. The comparison is conceptual, not a confirmed ranking factor or public Google score.

© Patrick Stox LLC · CC BY 4.0 ·

The Google patent behind the term

co it says

The patent jest “Contextual estimation of link information gain” (US20200349181A1, filed October 2018, granted June 2022; inventors obejmować Victor Carbune i Pedro Gonnet Anders). The core definition, verbatim z the patent:

“An information gain score for a given document is indicative of additional information that is included in the given document beyond information contained in other documents that were already presented to the user.”

It describes a maszyna-learning model że takes semantic vectors z documents the użytkownik już viewed plus data z a candidate new document, i outputs, in the patent’s words:

“a quantitative score between 0.00 and 1.00, with 0.00 indicating that no information gain is to be expected”

— z 1,00 meaning the document contains tylko information nie in co był już seen. Documents może then być ranked co najmniej partly on ich respective information gain scores.

co it doesn’t say

ten jest gdzie najbardziej SEO treść overreaches, so let me być precise o the limits:

  • No confirmed live używać. A granted patent jest dowód Google mógł deploy ten, nie że it robi. Google ma nie confirmed lub denied używając ten mechanism in ranking. Anyone telling you it’s a confirmed czynnik rankingowy jest guessing.
  • Narrower scope than assumed. The patent’s framing centers on choosing co document lub link to surface następny — think a conversational lub assistant follow-up po a użytkownik ma już viewed niektóre wyniki — nie on the pierwszy strona of ten blue links. Roger Montti’s analiza at wyszukiwarka Journal makes ten scope point well: the emphasis jest on automated assistants, i te scores aren’t described as applying to the pierwszy ustawić of wyniki (SEJ).
  • No public formula. The patent nazwy “semantic vectors” i salient extracted information, ale discloses no weights, no funkcja lista, i no reproducible scoring metoda. robić nie trust dowolny artykuł że hands you a precise “how Google calculates it” formula — it’s invented.
  • It’s a patent family, nie one filing. Google był granted a continuation in the same family — US12013887B2, same tytuł, same assignee (Google LLC), priority date back to the original October 2018 filing. A continuation means więcej disclosed twierdzenie język exists in the family, nie że anything new o deployment ma był confirmed — a granted patent i jego continuation nadal tylko establish co Google disclosed it mógł build, nie co’s running in production.

Bill Slawski’s early breakdown at Go Fish Digital framed the underlying problem Google jest solving — że gdy wiele documents share a topic, they tend to contain similar information — i how the patent proposes ranking partly on information gain scores to address it (Go Fish Digital).

Google’s official guidance że echoes the idea

Here’s the honest part najbardziej explainers skip: no Google document, blog post, lub spokesperson używa the phrase “information gain” in public search guidance. The term jest industry shorthand. ale Google’s rzeczywisty official guidance — “Creating helpful, reliable, people-first content” — repeatedly describes the same underlying idea in jego self-assessment questions. te są verbatim z że strona:

“Does the content provide original information, reporting, research, or analysis?”

“If the content draws on other sources, does it avoid simply copying or rewriting those sources, and instead provide substantial additional value and originality?”

“Does the content provide insightful analysis or interesting information that is beyond the obvious?”

“Does the content provide substantial value when compared to other pages in search results?”

że ostatni one — substantial wartość porównany to other strony in wyniki wyszukiwania — jest o as close as Google’s official język gets to the information-gain concept bez używając the term. ten jest the strongest legitimately-official tie-in available, i it’s why “information gain” caught on as użyteczny shorthand dla a rzeczywisty principle Google clearly cares o, even though it’s nie Google’s word.

robi Bing używać information gain?

nie że I’ve found. No Bing lub Microsoft document lub spokesperson używa the term in the material I’ve sprawdzony. Bing’s publicly discussed czynniki rankingowe — witryna i author reputation, treść completeness, semantic relevance — point in a similar direction (depth i originality versus peers), ale że’s an adjacent idea, nie the same twierdzenie i nie the same vocabulary. If someone cytaty Bing używając the literal phrase “information gain,” treat it z suspicion until you see the primary źródło.

Why information gain matters więcej in an AI-answer-heavy SERP

ten jest the part że makes information gain więcej niż a patent-trivia curiosity.

AI ma przeczytaj internet. ChatGPT, Claude, Gemini, i Google’s own Omówienia AI i AI Mode są excellent at one specific thing: synthesizing i repackaging co już exists w całym the web. Nathan Wahl at Animalz makes the argument sharply — te systemy może już reiterate i repackage everything że’s był opublikowany, so the question becomes why publikować anything że isn’t additive. gdy Google synthesizes an answer, it pulls z multiple źródła — Wahl’s framing jest że you don’t need to outrank giants if twój treść contains information theirs doesn’t (Animalz).

że’s the mechanism. If twój strona restates consensus, an AI Overview absorbs it i answers bez sending anyone to you. If twój strona contains something tylko you mieć — a proprietary liczba, a rzeczywisty experiment, an expert’s pierwszy-hand take — an AI system ma to draw on you specifically to obejmować że fakt, zamiast paraphrasing it z wherever else it appears. że’s nie a guarantee: pobieranie, inclusion in the model’s context, generation, i cytowanie są oddzielny kroki, i originality clearing one of them doesn’t mean it clears the rest. Scoped research on generative-engine visibility (Aggarwal et al., “GEO: Generative Engine Optimization”) finds że adding źródła, statistics, i quotations to treść może poprawić visibility in AI-generated answers in ich tested setups — dowód że originality pomaga, nie proof że it guarantees a cytowanie, a ranking, lub ruch. Information gain jest a precondition dla będąc cited zamiast paraphrased, nie a promise of it.

Bernard Huang at Clearscope ma tied ten to Google’s Knowledge Graph: aim dla treść że covers concepts i entities on the fringe of co Google już “knows” o a topic, ponieważ LLMs są great at consensus ale genuine novelty jest nadal gdzie humans dodawać wartość (Clearscope). Amanda King’s early wyszukiwarka Land framing captured the core idea — an information gain score jest essentially a mierzyć of how unique twój treść jest versus the rest of the corpus, i treść risks będąc demoted if it lacks uniqueness even gdy it’s just the same ideas in różny words (wyszukiwarka Land). Andrew Holland later reframed the whole thing dla the AI era: the term means różny things to różny people, ale the job jest to zachować increasing the rate of information gain — i increasingly we’re optimizing dla AI, nie just dla Google (wyszukiwarka Land).

I’ve seen ten z the other direction, too. I ran niektóre tests trying to rank treść in Google’s AI Mode, i the treść simply didn’t rank — even though it był, as far as I mógł tell, better i więcej relevant than strony że zrobił. One of the possibilities I floated publicly był a lack of information gain in the artykuły: if the treść jest a well-written restatement of co’s już out there, there może być nothing dla the system to reward (my post on X).

co actually produces information gain

Synthesizing w całym the źródła — i my own experience — here’s co genuinely moves the needle:

  • Original research, surveys, i experiments. The najbardziej defensible form. Run a study, publikować the data, let others cite you. At Ahrefs, nasz best-performing treść skews heavily toward data studies dla exactly ten powód.
  • Proprietary / pierwszy-party data. Internal produkt lub usage data nobody else może reproduce. If you mieć a dataset, że dataset jest twój information gain.
  • Expert interviews i pierwszy-hand experience. ten jest my go-to tactic. Rather than wysyłka unsourced generated tekst, message a handful of rzeczywisty subject-matter experts you już mieć access to — ponad Slack, email, lub a quick AI-assisted call — a kilka times a week, i capture ich rzeczywisty stories i experience. że’s original information że didn’t exist on the web przed you opublikowany it. I applied ten to my own rebuild of Ahrefs’ techniczne SEO hub (roughly 160 AI-generated strony) by layering expert sprawdzenie i czytelnik feedback on top so the strony carry something beyond regurgitated consensus.
  • Contrarian lub updated takes on consensus. Testing a widely repeated twierdzenie i reporting co actually happened jest information gain, even gdy the “study” jest mały.
  • Building on a predecessor’s działać. Take someone’s opublikowany research, extend it, dodaj następny data point. You’re adding to the corpus, nie copying it.

Ahrefs’ own treść-quality ladder (Si Quan Ong’s “How to Create Quality Content”) jest a użyteczny map here: it climbs z prosty listicles up to research studies i original ideas — effectively an information-gain ladder, prizing pierwszy-hand data ponad aggregation, bez ever używając the term.

How to evaluate twój own treść dla information gain

Adapt Google’s pomocny-treść questions do a pre-publikować gut sprawdzenie:

  • otwarty the top wyniki dla twój target zapytanie. lista co każdy one już says.
  • dla twój draft, highlight każdy zdanie że dodaje a fakt, liczba, angle, lub experience nie on że lista. If nothing jest highlighted, you mieć a rewrite, nie a zasób.
  • Ask whether an AI mógł answer the zapytanie fully z the strony już ranking. If yes, twój tylko path to relevance jest contributing something tamte strony lack.
  • Sanity-sprawdź “value compared to other pages in search results” question literally — nie “is this good,” ale “is this additive.”

The bottom wiersz

Information gain jest a rzeczywisty, użyteczny concept sitting on top of an unconfirmed mechanism. Don’t sprzedawać it internally as a confirmed Google czynnik rankingowy z a score you może dial — że’s nie prawdziwy, i it’ll burn twój credibility. sprzedawać it as the thing że’s demonstrably działający in an AI-heavy SERP: stop publishing better-optimized versions of co już exists, i start publishing co tylko you może. że’s the moat.

powiązany reading in ten cluster: entity SEO, schemat znaczniki dla AI, GEO, i AEO wszystkie connect to how you get surfaced i cited once twój treść actually ma something new to say.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.