Machine Translation and SEO
Is machine translation safe for SEO? Google's scaled content abuse policy targets bulk unreviewed MT for ranking manipulation — not translation itself. Here's the current policy, the Reddit case, the MTPE workflow, and the hreflang layer MT content still needs.
Machine translation isn't banned for SEO. What Google's scaled content abuse policy targets is publishing raw, unreviewed, bulk machine-translated pages at scale to manipulate rankings — the trigger is low value at scale, not the translation method. Google said as much in 2025 (AI-translated content is not 'strictly defined... as spam'), removed its old advice to block auto-translated pages with robots.txt, and took no action against Reddit's tens of millions of AI-translated URLs. The standard workflow at scale is MTPE — machine translation plus human post-editing — because 100% human translation of every page usually isn't practical. Raw unreviewed MT dumps are the actual risk. And translation quality is only half the job: my study of 374,756 hreflang-using domains found over 67% had some hreflang issue, so MT content frequently mis-targets its audience regardless of how good the translation is.
TL;DR — Machine translation — using a tool like Google Translate, DeepL, or an AI model to translate your pages — is not banned for SEO. What gets sites in trouble is publishing loads of raw, unchecked machine translations just to rank in more languages. If a human reviews and cleans up the translation before you publish, you’re doing what most big international sites already do.
What “machine translation and SEO” means
Machine translation (MT) is when software translates your content instead of a person — Google Translate, DeepL, Microsoft Translator, or an AI chatbot. The SEO question people actually ask is some version of: “If I use this to translate my site, will Google penalize me?”
The short answer: no, not for using machine translation. Google penalizes a pattern, not a tool.
The rule, in one sentence
Google’s guidelines target bulk, unreviewed, low-value pages published mainly to game rankings — and machine translation is just one way to churn those out. It sits in the same list as scraping and shuffling words around. The problem is “little value… to users” at scale, not the fact that a machine did the translating.
Evidence for this claim Google's scaled content abuse policy focuses on large amounts of low-value content created primarily to manipulate rankings and includes low-value automated translation as one possible technique. Scope: Google Search spam policy; machine translation is not categorically prohibited by this policy. Confidence: high · Verified: Google: Scaled content abuseSo there are really two very different things people call “machine translation”:
- A raw MT dump — you run a thousand pages through Google Translate and hit publish, nobody checks them. This is the risky one.
- MT + human review — the machine does the first draft, then a person fixes the awkward bits before it goes live. This is normal, and it’s fine.
That second workflow has a name — MTPE, machine translation post-editing — and it’s how most large multilingual sites operate. Translating every single page 100% by hand usually isn’t realistic once you have thousands of pages, so they lean on the machine and put humans on quality control.
Whatever the workflow, the visible main content—not just the surrounding boilerplate—needs to be coherently translated for the target audience.
Evidence for this claim Google recommends making a page's language obvious in visible content and warns that translating only boilerplate while leaving the main content unchanged can create a poor experience. Scope: Google Search guidance for multilingual page content. Confidence: high · Verified: Google: Make page language obviousWhat Google actually said (recently)
In 2025 Google said out loud that content translated by AI is not automatically spam, and it quietly deleted its old advice telling site owners to block auto-translated pages. It even left Reddit’s tens of millions of AI-translated pages alone. The direction is clear: Google judges the page by whether it’s useful, not by how it was made.
The catch beginners miss
Getting the translation right is only half the job. There’s a technical tag called hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. that tells Google which language version is for whom — and in my study of hundreds of thousands of websites, most of them had it broken. A perfectly translated page with broken hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. still gets shown to the wrong people. So “machine translation and SEO” is really two problems: is the translation good enough, and is the plumbing pointing it at the right audience.
Want the full version — the exact policy text, the Reddit case at scale, the real MTPE workflow, and how the hreflang layer breaks — plus what happens if you don’t translate at all? Switch to the Advanced tab.
Validate machine-translated pages before scaling
Source-language residue
Test to run: Scan rendered titles, headings, body text, navigation, structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding., and image text for unexpected source-language strings, then review flagged cases. Expected result: Only approved names and intentionally untranslated terms remain. Failure interpretation: The translation pipeline skipped fields or injected the wrong locale bundle. Monitoring window: Every publication batch and after template changes. Rollback trigger: Material source-language passages appear on production pages.
Human post-edit sample
Test to run: Have a qualified target-language reviewer score a representative sample for meaning, fluency, terminology, and local intent. Expected result: Pages communicate the source accurately and read naturally for the target audience. Failure interpretation: The model, glossary, or source content is unsuitable for unattended publication. Monitoring window: Before launch and for each materially new content type. Rollback trigger: Critical meaning, safety, legal, or brand errors escape the review sample.
Technical locale mapping
Test to run: Crawl translated pages and validate status, indexability, self-canonical, reciprocal hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others., and direct alternate targets. Expected result: Every approved translation is an accessible canonical member of the correct cluster. Failure interpretation: The publishing pipeline created content without a consistent technical mapping. Monitoring window: Immediately after each batch and through the next full crawl. Rollback trigger: Priority translations canonicalize to the source page or produce broken hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. clusters.
Post-launch value check
Test to run: Segment search impressions, clicks, engagement, and user feedback by translated directory and content type against its own prelaunch or prior-period baseline. Expected result: Approved pages earn relevant exposure without a concentrated quality complaint pattern. Failure interpretation: The batch may be low-value, poorly localized, or mapped to the wrong queries. Monitoring window: From indexation through the first seasonally comparable review period. Rollback trigger: A repeatable quality issue affects the batch and cannot be corrected page by page quickly.
Evidence for this claim Google's scaled content abuse policy focuses on large amounts of low-value content created primarily to manipulate rankings and includes low-value automated translation as one possible technique. Scope: Google Search spam policy; machine translation is not categorically prohibited by this policy. Confidence: high · Verified: Google: Scaled content abuse Evidence for this claim Google recommends making a page's language obvious in visible content and warns that translating only boilerplate while leaving the main content unchanged can create a poor experience. Scope: Google Search guidance for multilingual page content. Confidence: high · Verified: Google: Make page language obviousTL;DR — Machine translation isn’t banned. Google’s scaled content abuseScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. policy targets bulk, unreviewed MT published at scale to manipulate rankings — the trigger is “little value… to users,” not the method. Google stated in 2025 that AI-translated content is not “strictly defined… as spam,” removed its old robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere.-blocking advice, and took no action on Reddit’s tens of millions of AI-translated URLs. The scaled workflow is MTPE (MT + human post-editing), because 100% human translation of every page usually isn’t practical. Two separate failure modes: translation quality, and the technical wrapper — my study of 374,756 hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others.-using domains foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. 67%+ had some hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. issue, so MT content frequently mis-targets its audience even when the translation is fine. And if you don’t translate at all, Google may auto-translate your pages onto its own
translate.googsubdomain and keep the traffic.
Does Google penalize machine-translated content?
No — not for being machine-translated. This is the single most outdated claim in competitor content, most of which was written before Google’s March 2024 policy rename and its 2025 clarifications, so it still parrots a blanket “Google penalizes auto-translated content” line.
Here’s the operative policy. Google’s spam policies define scaled content abuse (the section renamed from “auto-generated content” in March 2024) as: “Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users.” Translation appears once, in a bullet list of examples, grouped with scraping and synonymizing: “Scraping feeds, search results, or other content to generate many pages (including through automated transformations like synonymizing, translating, or other obfuscation techniques), where little value is provided to users.”
Read that precisely. It doesn’t say “translated content” or “machine-translated content” is a violation. The load-bearing words are “to generate many pages” and “where little value is provided to users.” The pattern is bulk + low value + ranking-manipulation intent. Translation is named as one way to commit that pattern, in the same breath as scraping — not as a category of banned content.
Google’s 2025 statement — the clearest current line
The most quotable, most current confirmation came in June 2025. After Reddit scaled AI translations across its site, Glenn Gabe asked Google directly whether that was sanctioned, and a Google spokesperson responded ( reported by Search Engine Land): “While we don’t comment on the status of specific sites or pages, nor do we provide individualized support for any site, our policies do not strictly define content that has been translated by AI as spam. Our scaled content abuseScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. policy mentions automated transformations, including translations, as part of the overall warning against creating large amounts of unoriginal content that provides little to no value to users.”
That’s Google confirming, in plain language, exactly the distinction above: AI/machine translation is not “strictly defined… as spam”; the policy is about scaled, low-value output, not the translation method itself.
What changed in 2024–2025
Three concrete moves put the current policy in context:
- The March 2024 rename. “Auto-generated content” became scaled content abuse, reframing the whole area around value at scale rather than how content was created. The old “auto-generated content” help language (which literally listed “text translated by an automated tool without human review or curation before publishing” as an example) was folded into this value-based framing.
- The 2025 robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. guidance removal. Google
removed its longstanding advice
telling site owners to use
robots.txtto block auto-translated pages, characterizing it in the changelog as “This is a docs-only change, no change in behavior.” Once policy evaluates content by user value rather than creation method, a blanket “block all machine-translated pages” rule stopped being right. The correct tool for a specific low-quality translated page is now page-levelnoindex, not a sitewiderobots.txtblock. - Consistency over 15 years. None of this is really new. Google’s people (Mueller in 2010, Cutts in 2011) were drawing the same line — unreviewed automated translation vs. reviewed translation — long before “AI translation” was the term. It was never about translation as a technique; it was always about unreviewed automation at scale.
MT-then-human-review vs. raw bulk MT dumps
This is the distinction that actually matters operationally. Two things get called “machine translation for SEO,” and they live on opposite sides of the policy:
- Raw bulk MT dump — thousands of pages run through Google Translate or DeepL and published with zero human review, primarily to appear in more languages. This is the scaled-content-abuse risk.
- MTPE — machine translation post-editing — MT produces the first draft; a human linguist or fluent editor reviews and corrects it before publishing. Light MTPE fixes readability and obvious errors; full MTPE brings output up to human-translation quality.
MTPE is the de-facto industry-standard workflow at any real scale. It’s worth being precise about what it is: a risk-control practice, not an official Google compliance step. Google doesn’t grant translated pages a “reviewed” exemption or any other formal pass — human review just keeps the output on the right side of the value bar the scaled-content-abuse policy actually measures. The reason MTPE dominates is pragmatic: 100% human translation of every page usually isn’t practical or affordable once you’re maintaining thousands or millions of localized URLs across a dozen markets. Raw MT is the cheap-and-risky floor; full human translation is the expensive-and-slow ceiling; MTPE is where enterprise international SEOInternational SEO is the practice of optimizing a site so search engines understand which countries and/or languages it targets, and serve the right version to each user. It spans URL structure, hreflang, and on-page localization. actually lives. And doing all of it right — fluent output, human review, technically valid hreflang — still doesn’t guarantee indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., ranking, or how a page gets displayed; it removes the scaled-content-abuse risk, it doesn’t buy a ranking outcome. (I go deeper on the strategic layer — when you should translate at all versus fully localize — in the sibling piece on translation versus localizationLocalization is adapting content for a specific target market — not just translating the words, but adjusting currency, formats, idioms, cultural references, local search terms, and trust signals so the experience feels native..)
One more thing before you pipe anything through a translation tool: sending page content to an MT API or an LLMA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4). means that text leaves your system and lands with a third party. What happens to it — retention, whether it’s used for model training, where it’s stored, who else can access it — depends entirely on the specific provider, the plan you’re on, and your jurisdiction’s data-protection rules. That matters more for pages carrying personal data, regulated (health/financial/legal) content, or anything under an NDA or licensing restriction than for generic marketing copy. This isn’t SEO advice — check the provider’s current data-processing terms for your actual plan, and loop in whoever handles data protection/legal at your company, before you send anything sensitive through it. No SEO article, including this one, can clear that for you.
The Reddit case: proof it’s applied by value, not method
Reddit is the real-world stress test. Per Glenn Gabe’s reporting, Reddit scaled AI translations across 20+ languages and published tens of millions of AI-translated URLs — Gabe cites, for example, 2.3M ranking URLs in France and 2.4M in Spain. That’s about as “at scale” as scaled content gets. Google’s response, in Gabe’s words: “Well, nothing happened. Nothing at all.” No manual action, no algorithmic demotion.
The lesson isn’t “raw MT at scale is always safe” — it’s that Google applied the policy based on whether the underlying content was helpful, not on translation method or volume. (Be careful attributing the framing: Reddit characterizing the approach as “sanctioned” is Reddit’s framing, not a Google quote. What Google actually said is the “not strictly defined as spam” statement above.)
The technical layer MT content still needs
Here’s the differentiated point most coverage misses: translation quality and technical implementation are two separate failure modes, and you have to get both right.
Dedicated, indexable URLs — not JS overlays. A Google Translate widget or a client-side JS translation overlay does not give search engines real translated pages to rank. Google’s own localized-versions guidance and multi-regional docs assume distinct, crawlable URLs per language. On-the-fly translation without dedicated, indexable, hreflang-tagged URLs is a more clear-cut problem than MT quality — there’s simply nothing translated for the engine to index.
Hreflang — and why most of it is broken. hreflang tells search engines which language/region version to serve to whom. In my 374,756-domain hreflang study (Brighton SEO 2023), over 67% of domains using hreflang had at least one issue — missing x-defaultx-default is the reserved hreflang value that points to a fallback URL — the page Google shows when a user's locale doesn't match any of your other hreflang tags. It is optional and does not mean \"English.\" annotations, missing self-referencing tags, references to redirected or broken pages, missing reciprocal return tags, pointers to non-canonical URLs, and inconsistent language values. hreflang is one of the most complex aspects of SEO, and it only works as a reciprocal cluster: if two pages don’t both point at each other, Google ignores the pair entirely.
Put those together and the practical implication is uncomfortable: a majority of translated-page setups are mis-targeting their audience regardless of translation quality. A perfectly human-translated page with broken hreflang underperforms just as badly as a raw MT dump with perfect hreflang. Translation quality is necessary but not sufficient. (The mechanics — the three implementation methods, the reciprocity rule, x-default, valid codes — are in the hreflang and x-default deep dives.)
One useful clarification from Google’s docs: translated pages are not
automatically duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling.. Google explicitly treats a page as a duplicate only
“if the main content of the page remains untranslated” — i.e., the same
source-language body sitting on a /de/ URL. A genuinely translated page, however
it was produced, is not a duplicate.
The minimum gate, before anything else: a dedicated indexable URL per language, real main content visible without JavaScript, a reciprocal hreflangIf page A lists page B as an alternate, B must list A back — or Google ignores the pair. cluster (self-reference + every alternate + x-default), and a canonical that points at itself rather than the source-language page. Translation quality is worthless if any of those four are missing — Google has nothing indexable to judge. (The full walkthrough is in the Decision Trees lens.)
A render comparison can reveal when most visible content appears only after client-side JavaScript. It cannot judge translation quality, cultural fit, or whether a dedicated locale URL deserves indexing.
Run each localized template through my Render Gap Checker to compare its initial HTML with the rendered page before scaling publication. Render Gap Checker Free
- Test the public locale URL, not the source-language page with a browser widget switched on.
- Confirm the translated main content, links, canonical, and hreflang are present in dependable HTML.
- Send the language quality and market fit through a native-speaker MTPE review.
What happens if you don’t translate at all
There’s a live 2025 incentive to publish at least a reviewed baseline translation
rather than nothing: Google may auto-translate your content onto its own
translate.goog subdomain and keep the traffic. Per
Ahrefs’ analysis
(which I reviewed), an estimated 377M monthly organic visits flow through
Google’s translation-proxy pages, with India, Indonesia, and Brazil among the most
affected markets — traffic that could have gone to the original publisher’s own
localized pages.
My take, and I’ll stand behind it: Google has talked about improving the hreflang system for years; instead of continuing to help creators localize, it’s effectively decided to claim a chunk of that traffic as its own. Google frames the proxy as a fallback for when “there is no high-quality, local-language content available” (as Search Engine Land reported) — which flips the usual argument on its head. The absence of a real, reviewed translated page is what invites Google to translate it for you and keep the click. Shipping even a lean, MTPE-reviewed native-language page with proper hreflang is how you make your URL the one that gets indexed. This is a strong commercial argument for reviewed MT, not against translating.
Bing’s approach
Bing/Microsoft doesn’t publish a dedicated policy on machine translation the way
Google’s spam policy does. Its
Webmaster Guidelines
frame ranking around content quality and credibility broadly, without a
translation-specific carve-out. The honest read: Bing has no MT-specific rule, but
its general thin/low-quality-content guidance would apply the same way — unreviewed
bulk MT dumps are a subset of thin contentThin content is web content that provides little or no value to users. Google's spam policies name it 'thin content with little or no added value' — and it's about value per page, not word count., not a unique Bing policy. Note also that
Bing leans on the content-language signal more than hreflang, which is worth
accounting for in the technical wrapper.
The bottom line
Machine translation is a legitimate, mainstream SEO tool. Use it as a first draft, put a human on quality control (MTPE), give each language a dedicated indexable URL, get hreflang right, and judge the output by whether it’s actually helpful to that market. Do that and you’re not gaming anything — you’re doing what every large multilingual site already does. Skip the review and dump raw MT at scale to chase rankings, and that’s the pattern the scaled-content-abuse policy exists to catch.
AI summary
A condensed take on the Advanced version:
- MT is not banned. Google’s scaled content abuseScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. policy targets bulk, unreviewed, low-value pages published to manipulate rankings. Translation is named as one example alongside scraping/synonymizing; the trigger is “little value… to users,” not the method.
- Google, 2025: AI-translated content is not “strictly defined… as spam.” The scaled-content-abuse policy governs it by value.
- Two corroborating 2025 moves: Google removed its old robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. advice to
block auto-translated pages (“docs-only change, no change in behavior”), pointing
to page-level
noindexfor specific low-quality pages instead; and it took no action on Reddit’s tens of millions of AI-translated URLs. - The real distinction: raw bulk MT dump (risky) vs. MTPE — MT draft + human post-editing (standard practice). 100% human translation of every page usually isn’t practical at enterprise scale. MTPE is a risk-control practice, not an official Google compliance step — and even fluent output, human review, and valid hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. don’t guarantee indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., ranking, or display; they just remove the scaled-content-abuse risk.
- Two separate failure modes: translation quality and the technical wrapper. Patrick’s study of 374,756 hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others.-using domains foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. 67%+ had some hreflang issue, so MT content often mis-targets its audience regardless of translation quality.
- The minimum technical gate: dedicated indexable URL per language, real main content visible without JS, reciprocal hreflangIf page A lists page B as an alternate, B must list A back — or Google ignores the pair. + x-defaultx-default is the reserved hreflang value that points to a fallback URL — the page Google shows when a user's locale doesn't match any of your other hreflang tags. It is optional and does not mean \"English.\", self-canonical. Miss any of those and there’s nothing indexable for translation quality to matter.
- JS overlays aren’t translated pages. Search engines need dedicated, indexable, hreflang-tagged URLs per language.
- Before sending anything to an MT provider: check its data-handling/retention terms and loop in legal/privacy for sensitive or regulated content — that’s a data question, not an SEO one, and varies by provider, plan, and jurisdiction.
- Not duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling.: Google treats a page as a duplicate only if the main content remains untranslated.
- Don’t-translate risk: Google may auto-translate your pages onto
translate.goog(~377M monthly visits flow through it) and keep the traffic — an argument for reviewed MT. - Bing: no MT-specific policy; general thin-content guidance applies.
Official documentation
Primary-source documentation on machine-translated and localized content.
- Spam policies for Google web search — the scaled content abuseScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. section (renamed from “auto-generated content,” March 2024) names “automated transformations like synonymizing, translating, or other obfuscation techniques” where “little value is provided to users.” The judgment is value, not method.
- Using AI-generated content — Google’s helpful-content-first, method-agnostic stance, which the translation guidance mirrors.
- Localized versions of your pages — hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. mechanics; the reciprocity rule; and the note that a localized page is a duplicate “only if the main content of the page remains untranslated.”
- Managing multi-regional and multilingual sites — dedicated URLs per language, and the warning against translating only boilerplate. (This is the doc that previously carried the now-removed robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere.-blocking advice.)
- International SEO overview — locale-adaptive pages and how Google may not crawl/indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed./rank all locale variants.
Bing / Microsoft
- Bing Webmaster Guidelines — general content-quality expectations that unreviewed bulk MT can fall foul of; no MT-specific carve-out.
Quotes from the source
On-the-record statements from Google. Where a quote reached me through secondary reporting rather than a directly checkable source page, that’s noted.
Google — AI translation is not categorically spam (June 2025)
- “While we don’t comment on the status of specific sites or pages, nor do we provide individualized support for any site, our policies do not strictly define content that has been translated by AI as spam. Our scaled content abuse policyScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. mentions automated transformations, including translations, as part of the overall warning against creating large amounts of unoriginal content that provides little to no value to users.” — Google spokesperson, June 2025. Read the coverage
Google — the scaled content abuseScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. policy
- “Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users.” — Google Search Central docs. Jump to quote
- “…including through automated transformations like synonymizing, translating, or other obfuscation techniques…” — same page. Jump to quote
Google — the robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. guidance change (2025)
- “This is a docs-only change, no change in behavior.” — Google Search Central changelog, on removing the advice to block auto-translated pages via robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere.. Read the coverage
Google — localized pages and duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling.
- “Localized versions of a page are only considered duplicates if the main content of the page remains untranslated.” — Google Search Central docs. Jump to quote
Reddit case scale (via Glenn Gabe / GSQi)
- On Google’s (in)action against Reddit’s tens of millions of AI-translated URLs: “Well, nothing happened. Nothing at all.” — Glenn Gabe, GSQi. Read the coverage
Which translation method should this page get?
The question isn’t “MT or human?” in the abstract — it’s “how much human review does THIS page, in THIS market, justify?” Work it top to bottom.
1. Is this a dedicated, crawlable URL? If you’re only offering a Google Translate widget or a client-side JS overlay, stop — there’s no indexable translated page for search engines to rank. Create real, distinct URLs per language first. Everything below assumes you have them.
2. How much does this market matter?
- Low priority / tiny market / reference content with little nuance → raw MT is a defensible starting point, ideally clearly labeled and treated as a bridge.
- Any market you actually want to rank in → you need at least MTPE. Continue.
3. How competitive is the query / how high-value is the page?
- Non-competitive, informational, high volume of pages → light MTPE (MT draft
- human readability/error fixes). Good-enough beats perfect at scale.
- Mid-to-high value pages you want to compete on → full MTPE (MT draft + full human review to human quality).
- Highest-value money pages, brand-critical copy → human translation (or transcreation for taglines/ads).
4. Is it sensitive (checkout, legal, YMYL, lead-gen)?
Never ship raw MT here. Full MTPE or human, and consider an X-Robots-Tag to block
Google’s own proxy translation on those URLs.
5. Whatever tier you picked — is the wrapper right? Dedicated indexable URL ✓ · reciprocal hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. cluster (self + all alternates) ✓ · x-defaultx-default is the reserved hreflang value that points to a fallback URL — the page Google shows when a user's locale doesn't match any of your other hreflang tags. It is optional and does not mean \"English.\" set ✓ · canonical points to itself, not the source-language page ✓. Skip this and even a human-translated page mis-targets.
Rule of thumb: raw MT is a draft, not a publish-ready product for anything you want to rank. MTPE is the floor for competitive content. The scaled-content-abuse risk lives entirely in step 2 when you answer “low priority” for every market and ship the lot unreviewed.
Machine-translation myths that cost rankings
Each of these is a common belief from the wild, why it’s wrong, and what to do instead.
Myth: “Machine translation is banned / will get you penalized.” Why it’s wrong: Google explicitly stated AI/MT-translated content is not “strictly defined… as spam.” The risk is scaled, unreviewed, low-value output — not the translation method. Most content repeating this line predates Google’s 2024 policy rename. Do instead: Use MT freely as a draft; judge the published page by user value, and review before publishing.
Myth: “You must fully human-translate every page or Google will penalize you.” Why it’s wrong: MTPE (MT + human review) is standard practice at scale, and 100% human translation of every page usually isn’t practical for large sites. Google’s policy targets bulk unreviewed automation, not the presence of MT in the workflow. Do instead: Run MTPE — reserve full human translation for your highest-value pages and money/brand copy.
Myth: “A Google Translate widget or JS overlay = having translated pages.” Why it’s wrong: Client-side, on-the-fly translation gives search engines nothing indexable. There are no dedicated, crawlable, hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others.-tagged URLs to rank. Do instead: Publish real, distinct URLs per language with server-rendered translated content and reciprocal hreflangIf page A lists page B as an alternate, B must list A back — or Google ignores the pair..
Myth: “If Google auto-translates my page anyway, that covers my localizationLocalization is adapting content for a specific target market — not just translating the words, but adjusting currency, formats, idioms, cultural references, local search terms, and trust signals so the experience feels native..”
Why it’s wrong: Google’s translate.goog proxy is a fallback for when “there is
no high-quality, local-language content available” — it displaces your traffic and
branding onto Google’s own domain rather than fulfilling any strategy for you.
Do instead: Ship your own reviewed native-language page with hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. so your URL
is the indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. one; the proxy links tend to fall away once real hreflang’d pages
exist.
Myth: “This is a new problem created by AI.” Why it’s wrong: Google (Mueller in 2010, Cutts in 2011) drew the same unreviewed-vs-reviewed distinction long before “AI translation” was the phrase. The policy was never about the tool. Do instead: Treat it as the same old quality-at-scale question — reviewed automation is fine; unreviewed bulk dumps aren’t.
Myth: “Translated pages are duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling..” Why it’s wrong: Google treats a page as a duplicate only “if the main content of the page remains untranslated.” Different words in a different language aren’t duplicates. Do instead: Actually translate the main body (not just boilerplate), and mark variants with hreflang — don’t leave the source-language body on a country URL.
SOP: MTPE quality-control pass before publishing translated pages
A repeatable checklist your team runs on every batch of machine-translated pages before they go live. Adapt the depth (light vs. full MTPE) to the page’s value.
Prep
- Confirm the page has a dedicated, crawlable URL for the target language (not a JS overlay).
- Pull the source and MT output side by side in your review tool / TMS.
- Assign a reviewer who is native or fluent in the target language, not just a bilingual generalist.
Linguistic review (per page) 4. Read the MT output as a native user — flag anything that reads as machine output, awkward phrasing, or a mistranslated idiom. 5. Verify brand terms, product names, and UI strings are left untranslated or use the approved localized term (glossary/termbase). 6. Check numbers, currencies, units, dates, and legal/compliance phrasing — MT silently mangles these. 7. Confirm any text baked into images/screenshots is handled (MT won’t touch it).
SEO review (per page) 8. Confirm the title and meta descriptionThe meta description is an HTML head tag — `<meta name=\"description\" content=\"…\">` — that suggests a short summary of the page for the search snippet. It's not a Google ranking factor, and Google rewrites it the majority of the time, but a good one can still lift click-through. were translated and read naturally — not left in the source language. 9. Sanity-check the target keyword: is the MT’d phrasing what locals actually search, or a literal translation nobody uses? (This is the localizationLocalization is adapting content for a specific target market — not just translating the words, but adjusting currency, formats, idioms, cultural references, local search terms, and trust signals so the experience feels native. handoff, not just translation.) 10. Verify hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others.: self-reference present, all alternates listed, every alternate returns the reciprocal tag, x-defaultx-default is the reserved hreflang value that points to a fallback URL — the page Google shows when a user's locale doesn't match any of your other hreflang tags. It is optional and does not mean \"English.\" set. 11. Confirm canonical points to the page itself, not the source-language URL.
Publish & monitor
12. Publish; submit/refresh the sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. for that language.
13. After indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., check GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. for the language/country segment: is the right URL
ranking, or is a translate.goog proxy still showing? If the proxy persists,
re-audit hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others..
14. Log recurring MT error patterns back into the termbase / MT engine glossary so
the next batch’s raw output is better.
Cadence: run steps 4–11 on every page for high-value markets; sample (e.g., 10%) plus automated checks for low-priority, high-volume markets.
Playbook: you already published a raw MT dump — now what
A linear runbook for the situation where a bulk, unreviewed machine-translation rollout is already live and either underperforming or you’re worried about the scaled-content-abuse read. Work the steps in order.
Step 1 — Confirm the symptom.
Is it a quality/scale problem (thin, unreviewed pages at scale) or a targeting
problem (good pages, wrong audience)? Check GSCA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. by country/language: are your URLs
indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. and ranking, or are translate.goog proxies showing instead? The fix
differs.
Step 2 — Triage the pages, don’t panic-delete.
Segment by market value and traffic. You are not going to human-review a million
pages overnight, and you shouldn’t blanket-noindex everything either.
Step 3 — For genuinely low-value pages you can’t review soon:
Apply page-level noindex to the specific low-quality translated URLs (this is
Google’s post-2025 recommended tool — not a sitewide robots.txt block). Google’s
own remediation guidance for scaled content is to exclude it from Search if you’re
hosting it.
Step 4 — For pages in markets that matter: Run them through MTPE (see the SOP) starting with the highest-traffic/highest-value URLs. Fix the translation and the wrapper (hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others., canonical, titles/metas) in the same pass.
Step 5 — Fix the technical wrapper across the board. Even before linguistic review finishes, correct broken hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. clusters and self-canonical issues — this is often the bigger win, since 67%+ of hreflang-using domains in the study had some hreflang problem and mis-targeting suppresses even good pages.
Step 6 — Kill JS-only translation. If any “translated” content only existed as a Google Translate widget/overlay, replace it with real indexable URLs — there was nothing for engines to rank.
Step 7 — Re-index and verify.
Refresh sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. for the reviewed languages, request indexing on priority URLs, and
watch GSC by country: the goal is your URLs replacing any translate.goog proxy in
the SERP for the market.
Step 8 — Prevent recurrence. Move the MTPE SOP upstream so the next batch is reviewed before publishing. Feed recurring MT errors into the engine’s glossary/termbase.
Reassurance: Google took no manual action even against Reddit’s tens of millions of AI-translated URLs — the policy is applied on value, so a raw dump isn’t an automatic penalty. But “not penalized” isn’t “performing.” The playbook above is about making the pages actually work, which is the same thing that keeps them safe.
Real cases
Reddit — raw AI translation at massive scale, no action. Reddit scaled AI translations across 20+ languages and published tens of millions of AI-translated URLs (Glenn Gabe cites ~2.3M ranking URLs in France, ~2.4M in Spain). This is the biggest live test of the scaled-content-abuse policy against machine translation. Google’s response, per Gabe’s reporting: “Well, nothing happened. Nothing at all” — no manual action, no demotion. Google’s own statement (via Search Engine Land) was that AI-translated content is not “strictly defined… as spam.” Takeaway: the policy is applied on content value, not translation method or volume. (Attribute the “sanctioned” framing to Reddit, not Google.)
Google’s translate.goog proxy — the cost of not translating.
Ahrefs’ analysis
(which I reviewed) estimated 377M monthly organic visits flowing through Google’s
translation-proxy pages, with India, Indonesia, and Brazil among the hardest-hit
markets. When a publisher has no high-quality local-language page, Google translates
the English one onto its own translate.goog subdomain and keeps the click.
Before: no localized page → Google proxy captures the international traffic.
After: ship a reviewed native-language page with proper hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. → the proxy
links tend to disappear and your URL gets indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. for the market.
Untranslated boilerplate on a country URL — the actual duplicate-content trap.
A common failure: spin up /de/ URLs but leave the main body in English (only the
nav/footer translated). Google’s docs are explicit that a localized page is a
duplicate “only if the main content of the page remains untranslated.”
Before: English body on a German URL → treated as duplicate, not a real German
page. After: translate the main content (MT + review is fine), and Google has a
genuine German page to rank — no longer a duplicate.
Ready-to-use AI prompts
Copy-paste prompts for using an LLMA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4). in an MTPE workflow. Always keep a human in the loop — these speed up review, they don’t replace it.
Post-edit a raw machine translation (light MTPE)
You are a native [TARGET LANGUAGE] editor doing machine-translation post-editing.
Below is the [SOURCE LANGUAGE] original and a raw machine translation.
Fix the translation so it reads as if written by a native speaker: correct
awkward phrasing, mistranslated idioms, wrong register, and grammar. Do NOT
change meaning, do NOT translate brand/product names [LIST], and keep numbers,
currencies, dates, and units correct for [TARGET MARKET].
Return: (1) the corrected translation, and (2) a bullet list of every change you
made and why, so a human reviewer can spot-check.
SOURCE:
[paste]
RAW MACHINE TRANSLATION:
[paste]Flag likely mistranslations for human review (triage at scale)
Act as a QA reviewer for [TARGET LANGUAGE] machine-translated web content. Read
the translation below and output ONLY a table of suspected problems: the quoted
phrase, the issue type (idiom / mistranslation / wrong register / untranslated
term / number-format error / SEO keyword unnatural), and a suggested fix.
If nothing is wrong, say "no issues found." Do not rewrite the whole text.
TRANSLATION:
[paste]Check whether the translated keyword matches how locals actually search
For the [TARGET LANGUAGE / TARGET COUNTRY] market, is "[MACHINE-TRANSLATED
KEYWORD]" the phrase people actually search for this concept, or a literal
translation locals wouldn't use? Suggest 3-5 natural local alternatives and note
which is most likely to have search demand. Flag any that mean something
different locally (e.g., false-friend or regional-meaning traps).Localize the title and meta descriptionThe meta description is an HTML head tag — `<meta name=\"description\" content=\"…\">` — that suggests a short summary of the page for the search snippet. It's not a Google ranking factor, and Google rewrites it the majority of the time, but a good one can still lift click-through. (not just translate)
Translate and localize this page title and meta description for [TARGET
LANGUAGE / MARKET]. Keep the title under ~60 characters and the description under
~155. Use the natural local phrasing for the primary keyword rather than a literal
translation, and preserve the brand name [BRAND] untranslated.
TITLE: [paste]
META DESCRIPTION: [paste] Audit and detection snippets
Practical checks for finding raw-MT problems and verifying the technical wrapper on translated pages.
Detect a “translated” page that’s really a JS overlay
If the translated text only appears after JavaScript runs, search engines can’t indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. it. Compare raw HTML vs. rendered.
macOS / Linux (shell)
# Raw HTML the crawler sees first — does the translated body text appear here?
curl -sL "https://example.com/de/" | grep -o "EIN ERWARTETER DEUTSCHER SATZ"
# If that returns nothing but the text is visible in a browser, the translation
# is client-side only. Confirm with a real render (headless Chrome):
# npx -y @lighthouse ... or your renderer of choicePull the declared language and hreflang cluster from a URL
Chrome DevTools Console (paste on the page)
// Declared page language + every hreflang alternate on the page
console.table(
[...document.querySelectorAll('link[rel="alternate"][hreflang]')]
.map(l => ({ hreflang: l.hreflang, href: l.href }))
);
console.log('html lang =', document.documentElement.lang);
console.log('canonical =',
document.querySelector('link[rel="canonical"]')?.href);Check hreflang reciprocity across a set of URLs
hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. only works if every page in the cluster points back. This flags one-way tags — the single most common issue in my 374,756-domain study.
Python
import requests, re
from urllib.parse import urljoin
URLS = ["https://example.com/en/", "https://example.com/de/", "https://example.com/es/"]
def hreflangs(url):
html = requests.get(url, timeout=20).text
# crude but effective: grab rel=alternate hreflang link tags
tags = re.findall(
r'<link[^>]+rel=["\']alternate["\'][^>]+hreflang=["\']([^"\']+)["\'][^>]+href=["\']([^"\']+)["\']',
html, re.I)
return {lang: urljoin(url, href) for lang, href in tags}
clusters = {u: hreflangs(u) for u in URLS}
for u, alts in clusters.items():
for lang, target in alts.items():
back = clusters.get(target, {})
if u not in back.values():
print(f"NON-RECIPROCAL: {u} -> {target} ({lang}) has no return tag")Bookmarklet: highlight untranslated (source-language) blocks
Drop this as a bookmark; click it on a translated page to eyeball whether the main body was actually translated or left as boilerplate-only. Replace the word list with common source-language stopwords.
javascript:(()=>{const en=/\b(the|and|your|with|for|from|this)\b/gi;document.querySelectorAll('p,li,h1,h2,h3').forEach(el=>{const hits=(el.innerText.match(en)||[]).length;if(hits>=3)el.style.outline='2px solid red';});alert('Blocks outlined in red still look like source-language text.');})();These are diagnostics, not proof — always confirm findings by renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. the
page and reviewing with a native speaker. Adjust selectors/regex to your stack. Patrick's relevant free tools
- hreflang Generator + Linter — Enter your URL × locale matrix and get bidirectional hreflang markup as head tags, sitemap XML (auto-split past 50,000 URLs), and Link headers — linted live for wrong region codes, duplicates, and missing fallbacks. Runs entirely in your browser.
- returntag - hreflang checker — Enter one URL or an XML sitemap and the returntag - hreflang checker crawls the whole hreflang cluster — missing return tags, broken targets, self-reference and x-default checks, language-code validation, and head vs Link header vs sitemap disagreements — on an interactive cluster map with CSV export. The check Search Console's International Targeting report used to run.
- Local SERP Generator — Build a Google search link for a manual city or coordinate, country, language, and device context — without scraping results.
Tools for machine translation + SEO
Translation / MT engines
- DeepL — tends to edge out Google Translate on European language pairs; good raw quality for MTPE drafts.
- Google Translate / Cloud Translation API — widest language coverage; the 2025 Gemini integration improved idiom/context handling.
- Microsoft Translator — Azure-based, useful in Microsoft-stack workflows.
- LLMA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4). translation (Claude, GPT, Gemini) — strong for context-aware post-editing and glossary-constrained translation; pair with the Prompts lens.
Managing MTPE at scale
- TMS / localizationLocalization is adapting content for a specific target market — not just translating the words, but adjusting currency, formats, idioms, cultural references, local search terms, and trust signals so the experience feels native. platforms (e.g., Phrase, Crowdin, Lokalise, Smartling) — translation memory, termbases/glossaries, and human-review workflows so raw MT output improves over time.
Technical wrapper (hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. / indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.)
- Ahrefs Site Audit and Screaming Frog — crawl and flag hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. errors (missing self/return tags, non-canonical targets, broken alternates).
- Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. — International Targeting / performance by country to see
whether the right URL ranks per market (and whether a
translate.googproxy is showing instead). - hreflang tag generators / validators — build and check reciprocal clusters and x-defaultx-default is the reserved hreflang value that points to a fallback URL — the page Google shows when a user's locale doesn't match any of your other hreflang tags. It is optional and does not mean \"English.\" before shipping.
- Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility. — Bing leans on
content-language; verify indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. there separately.
Test yourself: Machine Translation and SEO
Five quick questions on how Google actually treats machine-translated content. Pick an answer for each, then check.
Resources worth your time
My related writing
- Google Is Stealing Your International Search Traffic With Automated Translations — the Ahrefs analysis I reviewed on Google’s
translate.googproxy pages capturing international traffic (~377M monthly visits) and how proper hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. wins it back. - The Beginner’s Guide to Technical SEO — where international and hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. mechanics fit in the bigger picture.
My speaking
- Hreflang Study and Interesting Issues (Brighton SEO 2023) — my 374,756-domain study; the 67%+-of-domains-have-hreflang-issues finding that underpins the “translation quality is only half the job” argument.
From around the industry
- Reddit uses AI to translate millions of pages and Google’s OK with it (Search Engine Land) — the June 2025 Google spokesperson statement in full.
- Is it safe? Google’s evolving view of auto-translated content (Glenn Gabe / GSQi) — the definitive deep dive on the Reddit case and the scaled-content-abuse framing.
- Google Removes Robots.txt Guidance For Blocking Auto-Translated Pages (Search Engine Journal) — the 2025 docs change and the shift toward page-level
noindex. - Is Google ‘stealing’ your international search traffic with translations? (Search Engine Land) — independent analysis of the proxy-translation phenomenon, with Google’s “no high-quality local-language content” framing.
- Spam policies for Google web search (Google) — read the scaled content abuseScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. section yourself; it’s shorter and clearer than most of the commentary about it.
- Localized versions of your pages (Google) — hreflang mechanics and the untranslated-main-content duplicate rule.
Machine Translation (for SEO)
Machine translation (MT) is using automated systems — Google Translate, DeepL, Microsoft Translator, or an LLM — to translate a site's content into other languages. It isn't banned for SEO; publishing raw, unreviewed MT in bulk purely to rank is what Google's scaled content abuse policy targets. MT reviewed and edited by a human (MTPE) is standard practice at scale.
Related: Localization, Hreflang, Duplicate Content
Machine Translation (for SEO)
Machine translation (MT) is the use of automated systems — Google Translate / Cloud Translation, DeepL, Microsoft Translator, or LLMA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4).-based translation — to produce non-source-language versions of a site’s content for search visibility in other markets. It is not, by itself, against Google’s or Bing’s guidelines.
What both search engines actually police is a subset of what Google now calls scaled content abuseScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. (renamed from “auto-generated content” in March 2024): publishing raw, unreviewed, bulk machine-translated pages at scale for the primary purpose of manipulating rankings, where little value is added for users. The policy names “translating” only as one example of an automated transformation alongside “synonymizing” and “scraping” — the trigger is low value at scale, not the translation method.
The de-facto industry-standard workflow is MTPE — machine translation post-editing — where MT produces the first draft and a human linguist reviews and corrects it. Light MTPE fixes readability and obvious errors; full MTPE brings the output up to human-translation quality. This is how most large multilingual and enterprise sites operate, because 100% human translation of every page usually isn’t practical or affordable at scale.
MT is distinct from localizationLocalization is adapting content for a specific target market — not just translating the words, but adjusting currency, formats, idioms, cultural references, local search terms, and trust signals so the experience feels native., which adapts the whole experience — local keywords, currency, formats, cultural references, and trust signals — for a market. MT only decides how the words get produced; it doesn’t localize them. And translation quality is separate from the technical wrapper: a well-translated page with broken hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. still mis-targets its audience, and a perfect hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. setup pointing at raw MT dumps still risks the scaled-content-abuse read. Both have to be right.
Related: Localization, Hreflang, Duplicate Content
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Updated Jul 18, 2026.
Editorial summary and recorded change details.Summary
Ran a research-informed accuracy pass: live-verified the current scaled-content-abuse policy wording (unchanged from what the article already quoted), reframed MTPE as a risk-control practice rather than an official Google compliance step, added an explicit note that fluent output/human review/valid hreflang don't guarantee indexing or ranking, surfaced the technical minimum gate directly in the Advanced technical-layer section, and added a non-legal-advice caution about checking a translation provider's data-handling terms before sending sensitive content.
Change details
-
Added a paragraph clarifying MTPE is a risk-control practice (not an official Google exemption) and that doing everything right doesn't guarantee ranking outcomes, in the Advanced 'MT-then-human-review' section.
-
Added a non-legal-advice caution about translation-provider data handling/retention and looping in legal/privacy for sensitive content.
-
Surfaced the technical minimum gate (dedicated URL, visible main content, reciprocal hreflang + x-default, self-canonical) directly in the Advanced technical-layer section, and echoed both additions in the AI Summary lens.
Full comparison unavailable — no prior snapshot was archived for this revision.