Machine Translation and SEO

Is machine translation safe for SEO? Google's scaled content abuse policy targets bulk unreviewed MT for ranking manipulation — not translation itself. Here's the current policy, the Reddit case, the MTPE workflow, and the hreflang layer MT content still needs.

First published: Jul 2, 2026 · Last updated: Jul 18, 2026 · Advanced
demand #4 in Localization & Content#16 in International SEO#381 on the site

Machine translation isn't banned for SEO. What Google's scaled content abuse policy targets is publishing raw, unreviewed, bulk machine-translated pages at scale to manipulate rankings — the trigger is low value at scale, not the translation method. Google said as much in 2025 (AI-translated content is not 'strictly defined... as spam'), removed its old advice to block auto-translated pages with robots.txt, and took no action against Reddit's tens of millions of AI-translated URLs. The standard workflow at scale is MTPE — machine translation plus human post-editing — because 100% human translation of every page usually isn't practical. Raw unreviewed MT dumps are the actual risk. And translation quality is only half the job: my study of 374,756 hreflang-using domains found over 67% had some hreflang issue, so MT content frequently mis-targets its audience regardless of how good the translation is.

TL;DR — Machine translation isn’t banned. Google’s scaled content abuseScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. policy targets bulk, unreviewed MT published at scale to manipulate rankings — the trigger is “little value… to users,” not the method. Google stated in 2025 that AI-translated content is not “strictly defined… as spam,” removed its old robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere.-blocking advice, and took no action on Reddit’s tens of millions of AI-translated URLs. The scaled workflow is MTPE (MT + human post-editing), because 100% human translation of every page usually isn’t practical. Two separate failure modes: translation quality, and the technical wrapper — my study of 374,756 hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others.-using domains foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. 67%+ had some hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. issue, so MT content frequently mis-targets its audience even when the translation is fine. And if you don’t translate at all, Google may auto-translate your pages onto its own translate.goog subdomain and keep the traffic.

Evidence for this claim Google's scaled content abuse policy focuses on large amounts of low-value content created primarily to manipulate rankings and includes low-value automated translation as one possible technique. Scope: Google Search spam policy; machine translation is not categorically prohibited by this policy. Confidence: high · Verified: Google: Scaled content abuse Evidence for this claim Google recommends making a page's language obvious in visible content and warns that translating only boilerplate while leaving the main content unchanged can create a poor experience. Scope: Google Search guidance for multilingual page content. Confidence: high · Verified: Google: Make page language obvious

Does Google penalize machine-translated content?

No — not for being machine-translated. This is the single most outdated claim in competitor content, most of which was written before Google’s March 2024 policy rename and its 2025 clarifications, so it still parrots a blanket “Google penalizes auto-translated content” line.

Here’s the operative policy. Google’s spam policies define scaled content abuse (the section renamed from “auto-generated content” in March 2024) as: “Scaled content abuse is when many pages are generated for the primary purpose of manipulating search rankings and not helping users.” Translation appears once, in a bullet list of examples, grouped with scraping and synonymizing: “Scraping feeds, search results, or other content to generate many pages (including through automated transformations like synonymizing, translating, or other obfuscation techniques), where little value is provided to users.”

Read that precisely. It doesn’t say “translated content” or “machine-translated content” is a violation. The load-bearing words are “to generate many pages” and “where little value is provided to users.” The pattern is bulk + low value + ranking-manipulation intent. Translation is named as one way to commit that pattern, in the same breath as scraping — not as a category of banned content.

Google’s 2025 statement — the clearest current line

The most quotable, most current confirmation came in June 2025. After Reddit scaled AI translations across its site, Glenn Gabe asked Google directly whether that was sanctioned, and a Google spokesperson responded ( reported by Search Engine Land): “While we don’t comment on the status of specific sites or pages, nor do we provide individualized support for any site, our policies do not strictly define content that has been translated by AI as spam. Our scaled content abuseScaled content abuse is Google's spam policy (introduced March 2024) for generating many low-value pages primarily to manipulate search rankings rather than help users — and it applies no matter how the content is created: AI, automation, or human writers. policy mentions automated transformations, including translations, as part of the overall warning against creating large amounts of unoriginal content that provides little to no value to users.”

That’s Google confirming, in plain language, exactly the distinction above: AI/machine translation is not “strictly defined… as spam”; the policy is about scaled, low-value output, not the translation method itself.

What changed in 2024–2025

Three concrete moves put the current policy in context:

  • The March 2024 rename. “Auto-generated content” became scaled content abuse, reframing the whole area around value at scale rather than how content was created. The old “auto-generated content” help language (which literally listed “text translated by an automated tool without human review or curation before publishing” as an example) was folded into this value-based framing.
  • The 2025 robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. guidance removal. Google removed its longstanding advice telling site owners to use robots.txt to block auto-translated pages, characterizing it in the changelog as “This is a docs-only change, no change in behavior.” Once policy evaluates content by user value rather than creation method, a blanket “block all machine-translated pages” rule stopped being right. The correct tool for a specific low-quality translated page is now page-level noindex, not a sitewide robots.txt block.
  • Consistency over 15 years. None of this is really new. Google’s people (Mueller in 2010, Cutts in 2011) were drawing the same line — unreviewed automated translation vs. reviewed translation — long before “AI translation” was the term. It was never about translation as a technique; it was always about unreviewed automation at scale.

MT-then-human-review vs. raw bulk MT dumps

This is the distinction that actually matters operationally. Two things get called “machine translation for SEO,” and they live on opposite sides of the policy:

  • Raw bulk MT dump — thousands of pages run through Google Translate or DeepL and published with zero human review, primarily to appear in more languages. This is the scaled-content-abuse risk.
  • MTPE — machine translation post-editing — MT produces the first draft; a human linguist or fluent editor reviews and corrects it before publishing. Light MTPE fixes readability and obvious errors; full MTPE brings output up to human-translation quality.

MTPE is the de-facto industry-standard workflow at any real scale. It’s worth being precise about what it is: a risk-control practice, not an official Google compliance step. Google doesn’t grant translated pages a “reviewed” exemption or any other formal pass — human review just keeps the output on the right side of the value bar the scaled-content-abuse policy actually measures. The reason MTPE dominates is pragmatic: 100% human translation of every page usually isn’t practical or affordable once you’re maintaining thousands or millions of localized URLs across a dozen markets. Raw MT is the cheap-and-risky floor; full human translation is the expensive-and-slow ceiling; MTPE is where enterprise international SEOInternational SEO is the practice of optimizing a site so search engines understand which countries and/or languages it targets, and serve the right version to each user. It spans URL structure, hreflang, and on-page localization. actually lives. And doing all of it right — fluent output, human review, technically valid hreflang — still doesn’t guarantee indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., ranking, or how a page gets displayed; it removes the scaled-content-abuse risk, it doesn’t buy a ranking outcome. (I go deeper on the strategic layer — when you should translate at all versus fully localize — in the sibling piece on translation versus localizationLocalization is adapting content for a specific target market — not just translating the words, but adjusting currency, formats, idioms, cultural references, local search terms, and trust signals so the experience feels native..)

One more thing before you pipe anything through a translation tool: sending page content to an MT API or an LLMA large language model (LLM) is a deep-learning model trained on massive text corpora to predict the next token and generate human-like text. LLMs use the transformer architecture and power AI search features like Google's AI Overviews (Gemini) and Bing Copilot (GPT-4). means that text leaves your system and lands with a third party. What happens to it — retention, whether it’s used for model training, where it’s stored, who else can access it — depends entirely on the specific provider, the plan you’re on, and your jurisdiction’s data-protection rules. That matters more for pages carrying personal data, regulated (health/financial/legal) content, or anything under an NDA or licensing restriction than for generic marketing copy. This isn’t SEO advice — check the provider’s current data-processing terms for your actual plan, and loop in whoever handles data protection/legal at your company, before you send anything sensitive through it. No SEO article, including this one, can clear that for you.

The Reddit case: proof it’s applied by value, not method

Reddit is the real-world stress test. Per Glenn Gabe’s reporting, Reddit scaled AI translations across 20+ languages and published tens of millions of AI-translated URLs — Gabe cites, for example, 2.3M ranking URLs in France and 2.4M in Spain. That’s about as “at scale” as scaled content gets. Google’s response, in Gabe’s words: “Well, nothing happened. Nothing at all.” No manual action, no algorithmic demotion.

The lesson isn’t “raw MT at scale is always safe” — it’s that Google applied the policy based on whether the underlying content was helpful, not on translation method or volume. (Be careful attributing the framing: Reddit characterizing the approach as “sanctioned” is Reddit’s framing, not a Google quote. What Google actually said is the “not strictly defined as spam” statement above.)

The technical layer MT content still needs

Here’s the differentiated point most coverage misses: translation quality and technical implementation are two separate failure modes, and you have to get both right.

Dedicated, indexable URLs — not JS overlays. A Google Translate widget or a client-side JS translation overlay does not give search engines real translated pages to rank. Google’s own localized-versions guidance and multi-regional docs assume distinct, crawlable URLs per language. On-the-fly translation without dedicated, indexable, hreflang-tagged URLs is a more clear-cut problem than MT quality — there’s simply nothing translated for the engine to index.

Hreflang — and why most of it is broken. hreflang tells search engines which language/region version to serve to whom. In my 374,756-domain hreflang study (Brighton SEO 2023), over 67% of domains using hreflang had at least one issue — missing x-defaultx-default is the reserved hreflang value that points to a fallback URL — the page Google shows when a user's locale doesn't match any of your other hreflang tags. It is optional and does not mean \"English.\" annotations, missing self-referencing tags, references to redirected or broken pages, missing reciprocal return tags, pointers to non-canonical URLs, and inconsistent language values. hreflang is one of the most complex aspects of SEO, and it only works as a reciprocal cluster: if two pages don’t both point at each other, Google ignores the pair entirely.

Put those together and the practical implication is uncomfortable: a majority of translated-page setups are mis-targeting their audience regardless of translation quality. A perfectly human-translated page with broken hreflang underperforms just as badly as a raw MT dump with perfect hreflang. Translation quality is necessary but not sufficient. (The mechanics — the three implementation methods, the reciprocity rule, x-default, valid codes — are in the hreflang and x-default deep dives.)

One useful clarification from Google’s docs: translated pages are not automatically duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling.. Google explicitly treats a page as a duplicate only “if the main content of the page remains untranslated” — i.e., the same source-language body sitting on a /de/ URL. A genuinely translated page, however it was produced, is not a duplicate.

The minimum gate, before anything else: a dedicated indexable URL per language, real main content visible without JavaScript, a reciprocal hreflangIf page A lists page B as an alternate, B must list A back — or Google ignores the pair. cluster (self-reference + every alternate + x-default), and a canonical that points at itself rather than the source-language page. Translation quality is worthless if any of those four are missing — Google has nothing indexable to judge. (The full walkthrough is in the Decision Trees lens.)

TIP Check whether the translated page exists before JavaScript runs

A render comparison can reveal when most visible content appears only after client-side JavaScript. It cannot judge translation quality, cultural fit, or whether a dedicated locale URL deserves indexing.

Run each localized template through my Render Gap Checker to compare its initial HTML with the rendered page before scaling publication. Render Gap Checker Free

  1. Test the public locale URL, not the source-language page with a browser widget switched on.
  2. Confirm the translated main content, links, canonical, and hreflang are present in dependable HTML.
  3. Send the language quality and market fit through a native-speaker MTPE review.
A client-dependent translation is a rendering risk even when the final browser view looks complete.

What happens if you don’t translate at all

There’s a live 2025 incentive to publish at least a reviewed baseline translation rather than nothing: Google may auto-translate your content onto its own translate.goog subdomain and keep the traffic. Per Ahrefs’ analysis (which I reviewed), an estimated 377M monthly organic visits flow through Google’s translation-proxy pages, with India, Indonesia, and Brazil among the most affected markets — traffic that could have gone to the original publisher’s own localized pages.

My take, and I’ll stand behind it: Google has talked about improving the hreflang system for years; instead of continuing to help creators localize, it’s effectively decided to claim a chunk of that traffic as its own. Google frames the proxy as a fallback for when “there is no high-quality, local-language content available” (as Search Engine Land reported) — which flips the usual argument on its head. The absence of a real, reviewed translated page is what invites Google to translate it for you and keep the click. Shipping even a lean, MTPE-reviewed native-language page with proper hreflang is how you make your URL the one that gets indexed. This is a strong commercial argument for reviewed MT, not against translating.

Bing’s approach

Bing/Microsoft doesn’t publish a dedicated policy on machine translation the way Google’s spam policy does. Its Webmaster Guidelines frame ranking around content quality and credibility broadly, without a translation-specific carve-out. The honest read: Bing has no MT-specific rule, but its general thin/low-quality-content guidance would apply the same way — unreviewed bulk MT dumps are a subset of thin contentThin content is web content that provides little or no value to users. Google's spam policies name it 'thin content with little or no added value' — and it's about value per page, not word count., not a unique Bing policy. Note also that Bing leans on the content-language signal more than hreflang, which is worth accounting for in the technical wrapper.

The bottom line

Machine translation is a legitimate, mainstream SEO tool. Use it as a first draft, put a human on quality control (MTPE), give each language a dedicated indexable URL, get hreflang right, and judge the output by whether it’s actually helpful to that market. Do that and you’re not gaming anything — you’re doing what every large multilingual site already does. Skip the review and dump raw MT at scale to chase rankings, and that’s the pattern the scaled-content-abuse policy exists to catch.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.