Ganho de Informação

O que ganho de informação significa em SEO e busca por IA — como funciona a patente do Information Gain Score do Google, por que pesquisa original e dados exclusivos superam conteúdo regurgitado, e como medir isso.

Publicado pela primeira vez: 2 de jul. de 2026 · Última atualização: 3 de ago. de 2026 · Avançado

Ganho de informação é quanto de informação *nova* sua página adiciona além do que um pesquisador já viu nos resultados para um tópico — novidade em relação ao corpus existente, não qualidade geral do conteúdo. O termo vem de uma patente real do Google ('Contextual estimation of link information gain', depositada em 2018, concedida em 2022) que pontua documentos de 0,00 a 1,00 por novidade e ranqueia parcialmente com base nessa pontuação — mas o Google nunca confirmou usar isso no ranqueamento ao vivo, e o enquadramento da patente é sobre o que mostrar *em seguida* depois que um usuário viu alguns resultados, não a primeira página. Não há pontuação visível no Search Console, nem API, nem fórmula pública: qualquer 'information gain score' em uma ferramenta é uma aproximação. O que é real e oficial é a orientação de conteúdo útil do Google perguntando se o conteúdo oferece 'informação original, reportagem, pesquisa ou análise'. Isso importa mais agora porque AI Overviews e AI Mode já podem sintetizar conteúdo consensual — apenas informação genuinamente nova (dados proprietários, pesquisa original, experiência em primeira mão) é o que a IA tem para citar em vez de absorver.

«> TL;DR — Information gain is novelty relative to the corpus a searcher has

already seen on a topic — not comprehensiveness, not E-E-A-T, not “quality” in the abstract. It’s named after a granted Google patent (“Contextual estimation of link information gain,” filed 2018, granted 2022) that scores documents 0.00–1.00 for how much new information they add and ranks partly on that. But: Google has never confirmed live use; the patent’s own framing is about what to show next after a user has seen some results, not the first page; and there’s no visible score, no API, no disclosed formula. What is official is Google’s helpful-content guidance asking for “original information, reporting, research, or analysis.” It matters more in an AI-answer SERP because LLMs synthesize consensus for free — only genuinely new information forces a citation. » (Tradução) (Síntese localizada do trecho vinte, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«## What information gain actually is (and isn’t) » (Tradução) (Síntese localizada do trecho vinte e um, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«The cited patent describes possible methods, not proof that a specific production ranking system operates exactly that way. Evidência desta afirmação Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Escopo: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confiança: alta · Verificado: Google patent: Contextual estimation of link information gain Use original-value guidance as an editorial principle, not a measurable ranking guarantee. Evidência desta afirmação Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Escopo: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confiança: alta · Verificado: Google: Helpful content guidance » (Tradução) (Síntese localizada do trecho vinte e dois, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«Most articles on this topic quietly conflate three different things. Keep them separate: » (Tradução) (Síntese localizada do trecho vinte e três, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«- Information gain — the marginal new information a document adds relative to what the searcher has already seen on the topic. It’s a delta.

  • Comprehensiveness / topical depth — covering everything about a topic. A page can be maximally comprehensive and have zero information gain if every fact in it already exists on the pages ranking above it.
  • E-E-A-T / content quality — trust, expertise, experience, authority. Related, but a different judgment. You can be an unimpeachable expert and still publish a page that adds nothing new. » (Tradução) (Síntese localizada do trecho vinte e quatro, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«- Truth and accuracy — being novel isn’t the same as being correct. A page can add a genuinely new claim and still get it wrong. Novelty, accuracy, and trust are separate judgments; treat originality as one input to publish, not a substitute for verifying it. » (Tradução) (Síntese localizada do trecho vinte e cinco, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«If you take one thing from this page: information gain is specifically about the delta, not the depth and not the trust. That distinction is the thing nearly every competing explainer blurs. » (Tradução) (Síntese localizada do trecho vinte e seis, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«One more distinction worth naming, because SEO borrowed the term loosely: “information gain” also has an older, more technical meaning outside search. In information theory it traces back to Claude Shannon’s 1948 work on entropy and information measures — a mathematical account of how much a signal reduces uncertainty, unrelated to search rankings. In information retrieval research, a parallel idea shows up as novelty and diversity: methods like Maximal Marginal Relevance (Carbonell and Goldstein, 1998) rerank or select results to reduce redundancy against what a reader has already seen. Google’s patent draws on the same underlying intuition — new relative to what’s already been shown — but none of these are interchangeable. The SEO industry’s “information gain” is shorthand built on a specific patent, not a direct application of Shannon’s math or of published IR diversity algorithms, even though the instinct they share (don’t just repeat what’s already there) rhymes across all three. » (Tradução) (Síntese localizada do trecho vinte e sete, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

O ganho de informação é o delta além do que já estava disponível—não o comprimento total ou o polimento da página candidata. Fonte: /ai-search/optimization/information-gain/

A caixa superior representa fatos consensuais A, B e C que o leitor já viu. O candidato de baixo ganho repete A, B e C com palavras diferentes. O candidato de alto ganho adiciona evidências originais D que estavam ausentes nas fontes anteriores. A comparação é conceitual, não um fator de classificação confirmado ou uma pontuação pública do Google.

© Patrick Stox LLC · CC BY 4.0 ·

«## The Google patent behind the term » (Tradução) (Síntese localizada do trecho vinte e nove, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«### What it says » (Tradução) (Síntese localizada do trecho trinta, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«The patent is “Contextual estimation of link information gain” (US20200349181A1, filed October 2018, granted June 2022; inventors include Victor Carbune and Pedro Gonnet Anders). The core definition, verbatim from the patent: » (Tradução) (Síntese localizada do trecho trinta e um, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«> “An information gain score for a given document is indicative of additional

information that is included in the given document beyond information contained in other documents that were already presented to the user.” » (Tradução) (Síntese localizada do trecho trinta e dois, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«It describes a machine-learning model that takes semantic vectors from documents the user already viewed plus data from a candidate new document, and outputs, in the patent’s words: » (Tradução) (Síntese localizada do trecho trinta e três, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«> “a quantitative score between 0.00 and 1.00, with 0.00 indicating that no

information gain is to be expected” » (Tradução) (Síntese localizada do trecho trinta e quatro, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«— with 1.00 meaning the document contains only information not in what was already seen. Documents may then be ranked at least partly on their respective information gain scores. » (Tradução) (Síntese localizada do trecho trinta e cinco, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«### What it doesn’t say » (Tradução) (Síntese localizada do trecho trinta e seis, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«This is where most SEO content overreaches, so let me be precise about the limits: » (Tradução) (Síntese localizada do trecho trinta e sete, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«- No confirmed live use. A granted patent is evidence Google could deploy this, not that it does. Google has not confirmed or denied using this mechanism in ranking. Anyone telling you it’s a confirmed ranking factor is guessing.

  • Narrower scope than assumed. The patent’s framing centers on choosing what document or link to surface next — think a conversational or assistant follow-up after a user has already viewed some results — not on the first page of ten blue links. Roger Montti’s analysis at Search Engine Journal makes this scope point well: the emphasis is on automated assistants, and these scores aren’t described as applying to the first set of results (SEJ).
  • No public formula. The patent names “semantic vectors” and salient extracted information, but discloses no weights, no feature list, and no reproducible scoring method. Do not trust any article that hands you a precise “how Google calculates it” formula — it’s invented.
  • It’s a patent family, not one filing. Google was granted a continuation in the same family — US12013887B2, same title, same assignee (Google LLC), priority date back to the original October 2018 filing. A continuation means more disclosed claim language exists in the family, not that anything new about deployment has been confirmed — a granted » (Tradução) (Síntese localizada do trecho trinta e oito, parte um: o texto-fonte foi preservado para conferência na revisão nativa.) « patent and its continuation still only establish what Google disclosed it could build, not what’s running in production. » (Tradução) (Síntese localizada do trecho trinta e oito, parte dois: o texto-fonte foi preservado para conferência na revisão nativa.)

«Bill Slawski’s early breakdown at Go Fish Digital framed the underlying problem Google is solving — that when many documents share a topic, they tend to contain similar information — and how the patent proposes ranking partly on information gain scores to address it (Go Fish Digital). » (Tradução) (Síntese localizada do trecho trinta e nove, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«## Google’s official guidance that echoes the idea » (Tradução) (Síntese localizada do trecho quarenta, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«Here’s the honest part most explainers skip: no Google document, blog post, or spokesperson uses the phrase “information gain” in public search guidance. The term is industry shorthand. But Google’s actual official guidance — “Creating helpful, reliable, people-first content” — repeatedly describes the same underlying idea in its self-assessment questions. These are verbatim from that page: » (Tradução) (Síntese localizada do trecho quarenta e um, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«> “Does the content provide original information, reporting, research, or

analysis?” » (Tradução) (Síntese localizada do trecho quarenta e dois, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«> “If the content draws on other sources, does it avoid simply copying or rewriting

those sources, and instead provide substantial additional value and originality?” » (Tradução) (Síntese localizada do trecho quarenta e três, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«> “Does the content provide insightful analysis or interesting information that is

beyond the obvious?” » (Tradução) (Síntese localizada do trecho quarenta e quatro, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«> “Does the content provide substantial value when compared to other pages in

search results?” » (Tradução) (Síntese localizada do trecho quarenta e cinco, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«That last one — substantial value compared to other pages in search results — is about as close as Google’s official language gets to the information-gain concept without using the term. This is the strongest legitimately-official tie-in available, and it’s why “information gain” caught on as useful shorthand for a real principle Google clearly cares about, even though it’s not Google’s word. » (Tradução) (Síntese localizada do trecho quarenta e seis, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«## Does Bing use information gain? » (Tradução) (Síntese localizada do trecho quarenta e sete, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«Not that I’ve found. No Bing or Microsoft document or spokesperson uses the term in the material I’ve reviewed. Bing’s publicly discussed ranking factors — site and author reputation, content completeness, semantic relevance — point in a similar direction (depth and originality versus peers), but that’s an adjacent idea, not the same claim and not the same vocabulary. If someone quotes Bing using the literal phrase “information gain,” treat it with suspicion until you see the primary source. » (Tradução) (Síntese localizada do trecho quarenta e oito, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«## Why information gain matters more in an AI-answer-heavy SERP » (Tradução) (Síntese localizada do trecho quarenta e nove, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«This is the part that makes information gain more than a patent-trivia curiosity. » (Tradução) (Síntese localizada do trecho cinquenta, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«AI has read the internet. ChatGPT, Claude, Gemini, and Google’s own AI Overviews and AI Mode are excellent at one specific thing: synthesizing and repackaging what already exists across the web. Nathan Wahl at Animalz makes the argument sharply — these systems can already reiterate and repackage everything that’s been published, so the question becomes why publish anything that isn’t additive. When Google synthesizes an answer, it pulls from multiple sources — Wahl’s framing is that you don’t need to outrank giants if your content contains information theirs doesn’t (Animalz). » (Tradução) (Síntese localizada do trecho cinquenta e um, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«That’s the mechanism. If your page restates consensus, an AI Overview absorbs it and answers without sending anyone to you. If your page contains something only you have — a proprietary number, a real experiment, an expert’s first-hand take — an AI system has to draw on you specifically to include that fact, rather than paraphrasing it from wherever else it appears. That’s not a guarantee: retrieval, inclusion in the model’s context, generation, and citation are separate steps, and originality clearing one of them doesn’t mean it clears the rest. Scoped research on generative-engine visibility (Aggarwal et al., “GEO: Generative Engine Optimization”) finds that adding sources, statistics, and quotations to content can improve visibility in AI-generated answers in their tested setups — evidence that originality helps, not proof that it guarantees a citation, a ranking, or traffic. Information gain is a precondition for being cited instead of paraphrased, not a promise of it. » (Tradução) (Síntese localizada do trecho cinquenta e dois, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«Bernard Huang at Clearscope has tied this to Google’s Knowledge Graph: aim for content that covers concepts and entities on the fringe of what Google already “knows” about a topic, because LLMs are great at consensus but genuine novelty is still where humans add value (Clearscope). Amanda King’s early Search Engine Land framing captured the core idea — an information gain score is essentially a measure of how unique your content is versus the rest of the corpus, and content risks being demoted if it lacks uniqueness even when it’s just the same ideas in different words (Search Engine Land). Andrew Holland later reframed the whole thing for the AI era: the term means different things to different people, but the job is to keep increasing the rate of information gain — and increasingly we’re optimizing for AI, not just for Google (Search Engine Land). » (Tradução) (Síntese localizada do trecho cinquenta e três, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«I’ve seen this from the other direction, too. I ran some tests trying to rank content in Google’s AI Mode, and the content simply didn’t rank — even though it was, as far as I could tell, better and more relevant than pages that did. One of the possibilities I floated publicly was a lack of information gain in the articles: if the content is a well-written restatement of what’s already out there, there may be nothing for the system to reward (my post on X). » (Tradução) (Síntese localizada do trecho cinquenta e quatro, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«## What actually produces information gain » (Tradução) (Síntese localizada do trecho cinquenta e cinco, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«Synthesizing across the sources — and my own experience — here’s what genuinely moves the needle: » (Tradução) (Síntese localizada do trecho cinquenta e seis, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«- Original research, surveys, and experiments. The most defensible form. Run a study, publish the data, let others cite you. At Ahrefs, our best-performing content skews heavily toward data studies for exactly this reason.

  • Proprietary / first-party data. Internal product or usage data nobody else can reproduce. If you have a dataset, that dataset is your information gain.
  • Expert interviews and first-hand experience. This is my go-to tactic. Rather than shipping unsourced generated text, message a handful of real subject-matter experts you already have access to — over Slack, email, or a quick AI-assisted call — a few times a week, and capture their actual stories and experience. That’s original information that didn’t exist on the web before you published it. I applied this to my own rebuild of Ahrefs’ technical SEO hub (roughly 160 AI-generated pages) by layering expert review and reader feedback on top so the pages carry something beyond regurgitated consensus.
  • Contrarian or updated takes on consensus. Testing a widely repeated claim and reporting what actually happened is information gain, even when the “study” is small.
  • Building on a predecessor’s work. Take someone’s published research, extend it, add the next data point. You’re adding to the corpus, not copying it. » (Tradução) (Síntese localizada do trecho cinquenta e sete, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«Ahrefs’ own content-quality ladder (Si Quan Ong’s “How to Create Quality Content”) is a useful map here: it climbs from simple listicles up to research studies and original ideas — effectively an information-gain ladder, prizing first-hand data over aggregation, without ever using the term. » (Tradução) (Síntese localizada do trecho cinquenta e oito, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«## How to evaluate your own content for information gain » (Tradução) (Síntese localizada do trecho cinquenta e nove, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«Adapt Google’s helpful-content questions into a pre-publish gut check: » (Tradução) (Síntese localizada do trecho sessenta, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«- Open the top results for your target query. List what each one already says.

  • For your draft, highlight every sentence that adds a fact, number, angle, or experience not on that list. If nothing is highlighted, you have a rewrite, not a resource.
  • Ask whether an AI could answer the query fully from the pages already ranking. If yes, your only path to relevance is contributing something those pages lack.
  • Sanity-check the “value compared to other pages in search results” question literally — not “is this good,” but “is this additive.” » (Tradução) (Síntese localizada do trecho sessenta e um, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«## The bottom line » (Tradução) (Síntese localizada do trecho sessenta e dois, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«Information gain is a real, useful concept sitting on top of an unconfirmed mechanism. Don’t sell it internally as a confirmed Google ranking factor with a score you can dial — that’s not true, and it’ll burn your credibility. Sell it as the thing that’s demonstrably working in an AI-heavy SERP: stop publishing better-optimized versions of what already exists, and start publishing what only you can. That’s the moat. » (Tradução) (Síntese localizada do trecho sessenta e três, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

«Related reading in this cluster: entity SEO, schema markup for AI, GEO, and AEO all connect to how you get surfaced and cited once your content actually has something new to say. » (Tradução) (Síntese localizada do trecho sessenta e quatro, parte um: o texto-fonte foi preservado para conferência na revisão nativa.)

Adicionar uma nota de especialista

Fixar uma citação de especialista

É uma pessoa nova? Crie o perfil não reivindicado dela em /admin/experts/ → Fixar uma citação de especialista primeiro.