Syndicated Content and SEO
When the same article runs on multiple sites, you pick one of three: cross-domain canonical, noindex the reprint, or risk a partner outranking you.
When the same article runs on multiple sites on purpose — wire pickups, partner reprints, cross-posting — you have to pick one of three things: (a) a cross-domain rel=canonical on the reprint pointing back to your original, (b) a noindex on the reprint, or (c) nothing, and let Google choose which copy to show. Google changed its guidance around May 2023: it no longer recommends canonical for syndication and now calls noindex 'the most effective solution,' because syndicated pages are 'often very different' and Google may not honor the canonical. Bing still prefers canonical where the partnership allows, plus syndicating excerpts not full text. Do nothing and a bigger partner can outrank you for your own story — whichever URL Google treats as canonical is the one that gets ranked. This is a policy decision, not a bug fix, and it's distinct from accidental duplicate content.
TL;DR — Syndicated contentSyndicated content is the deliberate, usually licensed republishing of the same article on other sites — wire pickups, partner reprints, or cross-posting. Because both parties know it's a copy, the SEO job is deciding which URL search engines should show, not eliminating a duplicate. is when the same article runs on more than one site on purpose — a wire service, a partner reprint, or you cross-posting your own post to Medium. When that happens you have to tell search engines which copy is the “real” one. You get three choices: put a canonical on the reprint pointing back to your original,
noindexthe reprint, or do nothing. Google now recommendsnoindex. Do nothing and a bigger partner can end up outranking you for your own story.
What syndicated content is
Syndication republishes content on another URL; Google may cluster duplicates and select a representative canonical. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Canonicalization Canonical declarations are signals rather than guaranteed controls over selection. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Canonicalization
Syndication is when you deliberately let the same article appear on other websites. Think of a newspaper story picked up by the Associated Press and run by dozens of member sites, a publisher licensing its articles to Yahoo or MSN, or a company reposting its own blog to Medium and LinkedIn for extra reach. Everyone involved knows it’s a copy — that’s the whole point.
That makes it different from duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling., which is usually an accident: the same page reachable through tracking parameters, a staging site that got indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., or someone scraping you without permission. With accidental duplication you try to get rid of the extra copy. With syndication both copies are supposed to exist, so the job isn’t “delete the duplicate” — it’s “tell search engines which one to show in results.”
The one decision you have to make
When the same story lives on several sites, search engines pick one of them to actually rank. If you don’t guide that choice, they’ll make it for you — and they can pick the bigger, more established partner site instead of you, the original author. That’s the risk: a partner outranking you for your own article.
You have three options:
- Canonical the reprint back to your original. The partner adds a line of
code (
rel="canonical") that says “the real version lives over there.” It’s a hint, though — search engines can ignore it. noindexthe reprint. The partner adds a tag that tells search engines “don’t put this copy in results at all.” This is the surest way to make sure your version is the one that shows.- Do nothing. You leave it to chance and hope Google shows your version. This is the risky path.
What Google recommends now
Google used to tell partners to add a canonical. It changed that advice around
May 2023. Now Google recommends noindex on the reprint as the most reliable
fix, because a syndicated article usually looks pretty different once it’s on the
partner’s site — different menus, ads, related links — and that difference makes
the canonical unreliable.
Bing is a little different: it still suggests the canonical approach where your agreement with the partner allows it, and it also likes the idea of syndicating just an excerpt with a link back, rather than the whole article.
The practical takeaway: don’t leave it to chance. Decide who should rank — you —
and make the partner either noindex their copy (Google’s pick) or canonical it
back to you. Want the exact tags, the reasoning, and how to check who’s actually
ranking? Switch to the Advanced tab.
TL;DR — Syndication is deliberate reuse, so the fix is a policy decision, not a bug patch. Three paths: (a) cross-domain
rel=canonicalon the reprint back to your original — Google’s former recommendation and still Bing’s stated preference, but fragile because syndicated pages are “often very different”; (b)noindexon the reprint (Googlebot-News: noindexfor News,Googlebot: noindex/robots: noindexfor Search) — Google’s current recommendation since ~May 2023; (c) do nothing and let Google pick, which can hand the ranking to a bigger partner. Whichever URL Google treats as canonical is the one rewarded, so leaving it ambiguous is the risky path. Bing additionally recommends excerpts over full text. Separately: syndication done mainly to build anchored backlinks at scale is a link-scheme violation regardless of canonical/noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed.. Publish-and-indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. the original first to reduce the “which is the source” ambiguity.
Syndication is not duplicate content — it’s a policy decision
Syndication is a distribution arrangement, while canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it. is Google’s duplicate-URL processing. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Canonicalization If exclusion of a partner copy is required, relying on rel=canonicalA rel=\"canonical\" annotation — in the HTML <head> or an HTTP Link header — that tells search engines which URL is the preferred version of duplicate or near-duplicate content. alone cannot guarantee it. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Canonicalization
Start here, because it reframes everything. Ordinary duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. is an accident — parameters, session IDs, a staging environment that leaked into the index, or somebody scraping you. The fix is to eliminate the extra copy or consolidate onto one URL.
Syndication is the opposite: it’s intentional, usually licensed, and both parties know the copy exists. The reprint is supposed to be there. So the SEO question isn’t “how do I remove this duplicate?” — it’s “which URL do I want search engines to rank, and how do I make that stick?” That’s a policy choice you make per syndication deal, not a bug you patch.
And it matters because whichever URL Google settles on as canonical is the one that gets the ranking benefit. As John Mueller put it when asked whether ranking signals consolidate onto a partner’s URL, “in general, if we recognize a page as canonical, that’s going to be the page most likely rewarded by our ranking systems.” If that page is your partner’s, your partner wins the search result for your story.
The core decision: three paths, one choice
Every syndication arrangement resolves to exactly one of these:
(a) Cross-domain rel=canonical on the reprint
The partner adds a canonical on their copy pointing at your original:
<link rel="canonical" href="https://www.your-original-site.com/the-article/">This was Google’s recommended approach for years. The theory: the canonical tells Google the partner’s page is a duplicate of yours and to consolidate onto yours. The problem — covered in detail below — is that a canonical tagA rel=\"canonical\" annotation — in the HTML <head> or an HTTP Link header — that tells search engines which URL is the preferred version of duplicate or near-duplicate content. is a hint, not a directive, and syndicated pages are frequently too different for Google to honor it. It’s still Bing’s stated preference where the partnership allows.
(b) noindex on the reprint
The partner adds a noindex so their copy is dropped from the index entirely:
<!-- Block regular Google Search (and everything else that honors robots) -->
<meta name="robots" content="noindex">
<!-- Block Google Search specifically -->
<meta name="googlebot" content="noindex">
<!-- Block Google News specifically -->
<meta name="googlebot-news" content="noindex">This is Google’s current recommendation. It doesn’t ask Google to agree the pages are duplicates — it removes the reprint from consideration altogether, so there’s no ambiguity left about which copy can rank. Your original is the only one eligible.
One precondition matters here: noindex only works if Google can actually crawl
the page and see the tag. Google’s own documentation is explicit — “for the
noindex rule to be effective, the page or resource must not be blocked by a
robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. file, and it has to be otherwise accessible to the crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index..” If the
partner’s robots.txt disallows the reprint’s URL (or the page requires a login,
or otherwise isn’t reachable), Google never sees the noindex and the reprint can
still surface via external links. Always check the partner’s robots.txt
alongside the tag itself.
(c) Do nothing
You leave both copies indexable and let Google’s canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it. pick a winner. The risk is concrete: Google may pick the partner. Glenn Gabe’s tracking study (more below) foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. no consistent logic in which copy Google surfaced when neither was blocked — including large partners duplicating or outranking the original in the News tabGoogle News SEO is the practice of getting eligible for and ranking well in Google's news-specific surfaces — the News tab of Search and the Top Stories carousel. As of the 2024–2025 Publisher Center transition there's no application to file: content that complies with Google's news content policies is automatically eligible, and ranking within that pool is driven by relevance, prominence, authoritativeness, freshness, usability, and location/language.. This is the option that gets a partner ranking for your own story.
Why canonical often fails for syndicated pages
This is the cruxChrome User Experience Report — Google's public dataset of real-world (field) performance data from eligible Chrome users. It's the official field-data source behind the Core Web Vitals program., and it’s the part most competing articles skip. A cross-domain
rel=canonical is built for pages that are near-identical. But once your article
lands in a partner’s CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms., it usually isn’t near-identical anymore: different
navigation, a different related-articles module, different ads, a different comment
section, sometimes a different headline or intro.
Google’s canonicalization logic clusters near-duplicate pages. When the syndicated version has enough surrounding differences, Google may decline to treat the two as duplicates for canonicalization at all — canonical tagA rel=\"canonical\" annotation — in the HTML <head> or an HTTP Link header — that tells search engines which URL is the preferred version of duplicate or near-duplicate content. or not — and rank them independently (or rank the wrong one). Google says exactly this in its troubleshooting doc: “The canonical link element is not recommended for those who want to avoid duplication by syndication partners, because the pages are often very different. The most effective solution is for partners to block indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. of your content.”
noindex sidesteps the whole problem. It doesn’t rely on Google agreeing the pages
are duplicates — it just removes the reprint from the index. That’s why Google moved
from a hint (canonical) to a directive (noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed.) for this specific case.
This ties directly into the site’s canonicalization framework: rel=canonical is a
hint, one of what Google’s Allan Scott has described as ~40 canonicalization
signals, and as I’ve framed it in that
deep dive, it’s a strong signal but “Google ignores it if other signals are
stronger.” Syndication is simply the scenario where that hint is least reliable —
which is exactly why the recommendation moved to a directive instead.
Google reversed its guidance around May 2023
Name this explicitly, because it’s the single most common piece of stale advice on the web right now — plenty of big marketing sites still lead with “just add a canonical.” Historically (including Google’s now-defunct 2009 “Handling legitimate cross-domain content duplication” blog post), the advice was to have partners add a cross-domain canonical. Around May 2023, Google updated its canonicalization troubleshooting doc to reverse that.
The dated anchor for the change is a John Mueller post from April 22, 2023 (relayed
via Search Engine Roundtable):
“For identical pages that you only slightly care which is picked, use
rel=canonical. For different pages (like syndication) and/or a strong opinion, use
noindex (+ maybe canonical).” Note that combination — noindex plus a
canonical is fine if you care more about keeping the reprint out of the index than
about consolidating signals.
Google’s own framing of why it changed, from Danny Sullivan (relayed via Search
Engine Journal):
“Our main help page change was to focus on your goal with syndicated contentSyndicated content is the deliberate, usually licensed republishing of the same article on other sites — wire pickups, partner reprints, or cross-posting. Because both parties know it's a copy, the SEO job is deciding which URL search engines should show, not eliminating a duplicate. rather
than the mechanism…” The goal — only one copy should be discoverable — matters more
than the specific tag, and noindex guarantees that goal better than a canonical
does.
Google’s guidance, in detail
Two docs carry the load:
- Fix canonicalization issues (troubleshooting) has a dedicated
#syndicated-contentsection — the source of the “not recommended… often very different… block indexing” statement above. It also lists syndicated content in its common-causes table and advises you to “advise syndication partners to block indexing.” - Avoid article duplication in Google News (News Publisher Center) is the doc
Google links to for the how. Its guidance: partners use
<meta name="Googlebot-News" content="noindex">to keep the copy out of Google News specifically, or the broader<meta name="Googlebot" content="noindex">(orrobotsnoindex) to block both News and regular Search. Its reasoning matches: the canonical is not recommended because “syndicated articles are often very different in overall content from original articles.”
So the mechanism is: pick the scope you need (News-only vs everything), and have the partner add the tag on their copy.
Bing’s guidance (don’t conflate the two engines)
Bing hasn’t followed Google here, so give it its own space. In its December 2025 post Does Duplicate Content Hurt SEO and AI Search Visibility?, Bing still recommends the canonical approach: “Ask partners to add a canonical tag pointing to your original URL when agreements allow.” It adds a tactic Google doesn’t emphasize — “When possible, syndicate excerpts instead of full articles, with a clear link back to the source.”
Bing also extends the concern to AI answers, not just classic rankings. Its framing: when “identical copies can exist across domains,” that “blurs intent signals,” making it harder for AI systems to decide which version best satisfies a query — which lowers the odds your preferred page gets used as a groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it. source. For 2026, that’s the AI-search hook: syndication ambiguity doesn’t just cost a blue link, it can cost you the citation in an AI answer.
When syndication crosses into a link scheme
There’s a separate, uglier version of “syndication” that’s really link-building in disguise: publishing the same (or spun) article across many sites mainly to plant keyword-anchored links back to your site. Google’s webspam guidance is explicit that “what does violate Google’s guidelines on link schemes is when the main intent is to build links in a large-scale way back to the author’s site” (relayed via Search Engine Land). Red flags Google named: keyword-rich anchor textAnchor text is the visible, clickable text of a hyperlink. It tells readers what they'll find on the other end and gives search engines context about the linked page. in syndicated bylines/bios, identical content across many publications, and partners not vetting author quality.
If that’s the situation, canonical/noindex isn’t enough — the fix is to nofollow
the outbound links (or canonical/noindex the whole page so it carries no ranking
credit). This is a different animal from legitimate wire-service or partner
reprints, and it’s worth separating so none of the above reads as an endorsement of
scaled syndication-for-links.
noindex doesn’t kill the backlink the way nofollow does
A common worry: “if the partner noindexes their copy, don’t I lose the link back
to me?” Not the same way a nofollow would. noindex removes the page from the
index; it doesn’t, by itself, strip a followed link on that page. Implemented as
noindex,follow, a link back to your original can still be followed. The honest
caveat: pages dropped from the index tend to get recrawled less often over time,
which can quietly reduce how much a link from them matters — but that’s a soft
effect, not the hard block that nofollow is. Don’t expect noindex to protect a
link forever, but don’t treat it as instantly killing one either.
Publish first, syndicate second
A practical tactic that reduces (doesn’t eliminate) the ambiguity: get your original crawled and indexed before — or at least at the same time as — partners publish theirs. Danny Sullivan framed the underlying problem as one of identification: “If people deliberately chose to syndicate their content, it makes it difficult to identify the originating source. That’s why we recommend the use of canonical or blocking.” Letting Google see your version first gives it a cleaner signal about which one is the origin. It won’t override a strong signal from a much larger partner, but it stacks in your favor.
How to monitor and recover
Watch for a partner outranking you: run manual SERP checks for your headline in
quotes, and check Search Console — URL Inspection to see
your Google-selected canonical, and the Performance reportThe Google Search Console report that shows how your site actually performed in Google Search, built from real impressions and clicks. It reports four metrics — clicks, impressions, average CTR, and average position — and keeps the most recent 16 months of data. to catch a story that
underperforms where a partner is visible instead. If a partner is already
outranking you, the escalation is: ask them to add noindex (or a canonical back to
you) per your agreement; if they won’t, that’s a contract/relationship problem more
than a technical one. Build the tag requirement into the syndication agreement up
front so it isn’t a negotiation after the fact.
AI summary
A condensed take on the Advanced version:
- Syndication is deliberate, licensed reuse — wire pickups, partner reprints, cross-posting. It’s not accidental duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. (parameters, staging, scraping). The fix is a policy decision — “which URL should rank?” — not a bug patch.
- Three paths, pick one: (a) cross-domain
rel=canonicalon the reprint back to your original; (b)noindexon the reprint; (c) do nothing and let Google pick. - Google recommends (b) noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. since ~May 2023, reversing its old canonical
advice. Its doc: canonical is “not recommended… because the pages are often very
different. The most effective solution is for partners to block indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed..” News
Publisher Center specifies
Googlebot-News: noindex(News only) orGooglebot: noindex/robots: noindex(News + Search). - Why canonical fails here:
rel=canonicalis a hint built for near-identical pages; syndicated pages differ in chrome/ads/nav, so Google may not treat them as duplicates at all.noindexis a directive that removes the copy outright. noindexhas a precondition: it only works if Google can crawl the page and see the tag. If the partner’srobots.txtblocks the reprint (or it’s otherwise unreachable), Google never sees thenoindexand the copy can still surface via external links.- Do nothing = the risk. Whichever URL Google treats as canonical gets the ranking benefit; that can be the bigger partner. Glenn Gabe’s 3,000-URL study foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. “no rhyme or reason” in which copy surfaced when neither was blocked.
- Bing differs: still prefers the canonical where agreements allow, and recommends syndicating excerpts not full articles. It also warns duplicate copies “blur intent signals” for AI-search groundingGrounding is anchoring an AI model's answer to source documents it retrieves at the moment you ask — not to the patterns frozen into its weights during training. Retrieval-Augmented Generation (RAG) is the most common way to do it..
- Link-scheme line: syndicating mainly to build anchored backlinks at scale
violates Google’s link-scheme guidance regardless of canonical/noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. —
nofollowthe links or noindex the whole page. noindex≠nofollow:noindex,followcan still pass a link back; noindexed pages just tend to get recrawled less over time. It’s not a hard link block.- Publish-and-index first, then syndicate, to give Google a cleaner “which is the source” signal.
Which path: canonical, noindex, or nothing?
Answer a few questions to land on the right instruction to give your syndication partner.
Official documentation
Primary-source guidance from the search engines.
- Fix canonicalization issues (troubleshooting) — the
#syndicated-contentsection: canonical is “not recommended” for syndication, “the most effective solution is for partners to block indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed..” This is Google’s current, load-bearing position. - Avoid article duplication in Google News (News Publisher Center) — the “how”:
Googlebot-News: noindex(News) vsGooglebot: noindex(News + Search), and the “often very different in overall content” reasoning. - How to specify a canonical URL with rel=canonical — the mechanics of the canonical method, for the path where you (or Bing) still want to use it.
- Block search indexing with noindex — how
noindexworks, the meta tag vs X-Robots-Tag headerThe X-Robots-Tag is an HTTP response header that carries the same indexing and serving directives as the robots meta tag (noindex, nofollow, nosnippet, and the rest). Because it lives in the header rather than the HTML, it's how you control indexing for non-HTML files like PDFs, images, and videos., and thenoindex,followcombination. - Link spam / link schemes (spam policies) — where scaled syndication-for-links crosses the line.
Bing / Microsoft
- Does Duplicate Content Hurt SEO and AI Search Visibility? (Dec 2025) — Bing’s current take: canonical where agreements allow, syndicate excerpts, and the AI-search “intent signals” angle.
Quotes from the source
On-the-record statements from Google and Bing, plus relayed rep statements. Each deep link jumps to the quoted passage where the source page supports it.
Google — canonical is no longer recommended for syndication
- “The canonical link elementA rel=\"canonical\" annotation — in the HTML <head> or an HTTP Link header — that tells search engines which URL is the preferred version of duplicate or near-duplicate content. is not recommended for those who want to avoid duplication by syndication partners, because the pages are often very different. The most effective solution is for partners to block indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. of your content.” — Google Search Central, Fix canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it. issues. Jump to quote
- “…include incorrect language annotations, faulty CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms. settings, server misconfigurations, malicious hacks, syndicated contentSyndicated content is the deliberate, usually licensed republishing of the same article on other sites — wire pickups, partner reprints, or cross-posting. Because both parties know it's a copy, the SEO job is deciding which URL search engines should show, not eliminating a duplicate., and copycat websites.” — Google Search Central, Fix canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it. issues (common causes). Jump to quote
John Mueller, Google — when to use noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. vs canonical (April 22, 2023)
- “For identical pages that you only slightly care which is picked, use rel=canonical. For different pages (like syndication) and/or a strong opinion, use noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. (+ maybe canonical).” (Relayed via Search Engine Roundtable; posted to X on April 22, 2023 — the dated anchor for the guidance change.) Read the coverage
Danny Sullivan, Google — why the guidance changed (2023)
- “Our main help page change was to focus on your goal with syndicated content rather than the mechanism…” (Relayed via Search Engine Journal.) Read the coverage
Danny Sullivan, Google — the identification problem (Sep 2019)
- “If people deliberately chose to syndicate their content, it makes it difficult to identify the originating source. That’s why we recommend the use of canonical or blocking. The publishers syndicating can require this.” (Relayed via Search Engine Land.) Read the coverage
John Mueller, Google — whichever URL is canonical gets rewarded
- “It’s complicated, and not all the things you’re asking about are things we necessarily even use. In general, if we recognize a page as canonical, that’s going to be the page most likely rewarded by our ranking systems.” (Relayed via Search Engine Journal, in a thread with Lily Ray.) Read the coverage
Google — when syndication becomes a link scheme
- “…what does violate Google’s guidelines on link schemes is when the main intent is to build links in a large-scale way back to the author’s site.” (Relayed via Search Engine Land.) Read the coverage
Microsoft Bing — canonical, excerpts, and AI intent (Dec 2025)
- “When your articles are republished on other sites, identical copies can exist across domains, making it harder for search engines and AI systems to identify the original source.”
- “Ask partners to add a canonical tagA rel=\"canonical\" annotation — in the HTML <head> or an HTTP Link header — that tells search engines which URL is the preferred version of duplicate or near-duplicate content. pointing to your original URL when agreements allow.”
- “When possible, syndicate excerpts instead of full articles, with a clear link back to the source.” — Microsoft Bing Blogs. Read the post
Runbook: a partner is outranking you for your own story
Linear steps for the incident where you find a syndication partner ranking above (or instead of) your original.
- Confirm it’s actually happening. Search your exact headline in quotes on Google (and the News/Top Stories tab if it’s a news piece). Note which URL shows — yours, the partner’s, or both. Do the same on Bing if it’s a priority.
- Check Google’s canonical choice. In Search Console → URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version. on your original, read Google-selected canonicalA Google Search Console Page Indexing status: you declared a canonical for this URL, but Google overrode your choice, picked a different page as the canonical, and indexed that one instead.. If Google selected the partner’s URL (or a “Duplicate, Google chose a different canonical” state), that confirms the diagnosis.
- Confirm the reprint’s current setup. Fetch the partner’s copy and check
whether it has a
noindex, a canonical pointing at you, or neither (curl -s <partner-url> | grep -iE 'noindex|canonical'). “Neither” is the usual cause. - Pull the agreement. Check whether your syndication contract already requires
the partner to
noindexor canonical. If it does, you’re enforcing a term, not asking a favor. - Request the fix. Ask the partner to add
noindexon their copy (Google’s preference) —<meta name="robots" content="noindex">for Search, orGooglebot-Newsscope if it’s only a News-tab problem. If Bing matters too, ask fornoindex+ a canonical back to you. - Verify the tag went live — and that Google can see it. Re-fetch the partner
URL and confirm the tag is in the rendered
<head>, not just promised. CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms. templates strip these more often than you’d think. Also check the partner’srobots.txt:noindexonly takes effect if the page isn’t blocked there and is otherwise reachable to the crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. — a blocked or gated page means Google never sees the tag. - Wait for a recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial..
noindexonly takes effect once Google recrawls the reprint. It’s not instant — the partner’s copy has to be re-fetched and dropped. Don’t expect same-day movement. - Recheck rankings and canonical selectionHow search engines pick one canonical URL among duplicates and consolidate signals onto it. after Google has had time to recrawl both URLs. Your original should reclaim the result once the reprint is out of the indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed..
- If the partner won’t comply, it’s a relationship/contract problem, not a technical one — escalate through the business relationship, and build the requirement into the next agreement so it’s non-negotiable up front.
SOP: setting up a new syndication deal (do this every time)
A repeatable checklist to run whenever you sign or renew a syndication arrangement, so attribution is settled before the content goes live — not after a partner is already outranking you.
- Decide who should rank. Default: your original. If you’re syndicating purely for referral traffic/brand and genuinely don’t mind the partner ranking, that’s a deliberate exception — write it down.
- Pick the instruction using the Decision Tree:
noindex(Google’s pick), canonical (Bing / where you prefer the hint), ornoindex+ canonical (both engines). - Put the tag requirement in the contract. Specify the exact tag, the scope (News-only vs everything), and that it must be present at publish time. This is the single highest-leverage step — it turns a later favor into an enforceable term.
- Sequence the publishing. Publish your original and get it indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. first (or simultaneously), so Google sees your version as the source. Confirm it’s indexed in Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. before the partner goes live if timing allows.
- Verify the reprint on launch day. Fetch the partner’s live URL and confirm
the agreed tag is present in the rendered
<head>— and confirm the URL isn’t blocked in the partner’srobots.txt, since a blocked page means Google can’t see thenoindexat all. - Schedule a recheck ~2–4 weeks out: SERP check for the headline + Search Console canonical selectionHow search engines pick one canonical URL among duplicates and consolidate signals onto it., to confirm your original is the version ranking.
- Document the deal (partner, URLs, chosen instruction, date) so the next audit knows the intended state and can spot drift if a CMSA content management system (CMS) is software that lets users create, manage, and publish digital content — like blog posts and pages — without writing raw code. WordPress, Drupal, and Joomla are the most common open-source CMS platforms. change strips the tag.
Common mistakes (and what to do instead)
“Just add a canonical tagA rel=\"canonical\" annotation — in the HTML <head> or an HTTP Link header — that tells search engines which URL is the preferred version of duplicate or near-duplicate content. and you’re covered.”
Why it’s wrong: Google specifically stopped recommending canonical for
syndication in ~May 2023. A canonical is a hint, and syndicated pages are “often
very different,” so Google frequently won’t honor it.
Do instead: Have the partner noindex the reprint — Google’s stated most-effective
solution. Add a canonical too if Bing matters, but don’t rely on canonical alone for
Google.
“Syndicated contentSyndicated content is the deliberate, usually licensed republishing of the same article on other sites — wire pickups, partner reprints, or cross-posting. Because both parties know it's a copy, the SEO job is deciding which URL search engines should show, not eliminating a duplicate. gets a duplicate-content penalty.” Why it’s wrong: There’s no blanket duplicate-content penalty (see duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling.). The real risk is losing visibility/attribution to the wrong copy — or, separately, a link-scheme violation if you syndicate mainly for backlinks. Do instead: Frame it as an attribution decision, not a penalty to avoid. Pick canonical/noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. based on who you want to rank.
“Google always shows the original publisher.”
Why it’s wrong: It doesn’t. Glenn Gabe’s 3,000-URL study foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. “no rhyme or reason”
in which copy Google surfaced when neither was blocked — including large partners
duplicating or outranking the original in the News tabGoogle News SEO is the practice of getting eligible for and ranking well in Google's news-specific surfaces — the News tab of Search and the Top Stories carousel. As of the 2024–2025 Publisher Center transition there's no application to file: content that complies with Google's news content policies is automatically eligible, and ranking within that pool is driven by relevance, prominence, authoritativeness, freshness, usability, and location/language..
Do instead: Don’t leave it to chance. Use noindex (or canonical) so the choice
isn’t Google’s to make.
“noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. will kill the backlink value the partner sends me.”
Why it’s wrong: noindex isn’t nofollow. Implemented as noindex,follow, a link
back can still be followed; noindex removes the page from the indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., it doesn’t
strip the link the way nofollow does.
Do instead: Use noindex,follow if you want the copy out of the index but the link
preserved. Just know noindexed pages get recrawled less over time, so don’t treat the
link as permanent.
“Syndication is the same as scraping / content theft.” Why it’s wrong: Syndication is licensed and agreed; scraping is unauthorized. Google treats scraped, no-value auto-content as a spam violation — a categorically different thing from a wire-service or partner deal. Do instead: Handle real syndication with canonical/noindex; handle scrapers with DMCA/enforcement, not a canonical request they’ll never honor.
“More syndication = more SEO value, so syndicate everywhere.” Why it’s wrong: If the same article lives on five domains, typically only one surfaces in search or AI answers — the rest are wasted from a rankings standpoint. Do instead: Syndicate for referral traffic and brand reach (legitimate reasons), but don’t expect extra ranking value from extra copies. Protect the one URL you want to rank.
Real scenarios, before and after
1. Wire pickup outranking the origin (news)
Before: A regional publisher’s original story gets picked up by a wire and run,
unchanged and indexable, on a much larger national partner. Neither copy is blocked.
Google surfaces the national partner in Top Stories; the origin is buried. This is
the exact pattern Glenn Gabe documented across ~3,000 URLs — “no rhyme or reason,”
with large partners like Yahoo duplicating or outranking the original in the News
tab.
After: Partners add <meta name="googlebot-news" content="noindex"> (or full
noindex) per the syndication terms. The reprints drop out, and the origin reclaims
Top Stories for its own reporting.
2. Marketing cross-post to Medium eating the blog post Before: A B2B team reposts its full blog article to Medium for reach. Medium’s domain authority is high, and the Medium copy starts ranking for the target query instead of the company’s own blog. The company is now competing with its own content on a platform it doesn’t control. After: They set the Medium copy to canonical back to the original (Medium supports a canonical import), and going forward syndicate an excerpt with a link back rather than the full text — the tactic Bing explicitly recommends. The company blog holds the ranking; Medium drives referral traffic.
3. “Syndication” that’s really a link scheme
Before: An agency places the same 800-word article, with a keyword-rich anchor link
in the author bio, across 40 low-vetting industry sites — purely to build links. This
matches Google’s described link-scheme pattern (“main intent is to build links in a
large-scale way”).
After: The links are nofollowed (or the pages noindexed) so they carry no ranking
credit, and the program is re-scoped toward genuine editorial placements. The
canonical/noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. question is secondary here — the link intent is the actual problem.
4. Doing nothing on purpose (a valid choice)
Before: A small site syndicates a how-to to a large partner mainly for the referral
traffic and brand exposure, and doesn’t block anything. The partner outranks them for
the topic.
After: They confirm this is fine for this piece — the goal was reach, not
rankings — and document it as a deliberate exception. Not every syndication needs a
noindex; the mistake is doing nothing by default rather than by decision.
Syndication SEO checklist
Run this before and after any syndication deal goes live:
- Confirmed this is real syndication, not scaled syndication-for-links (if
the latter,
nofollowthe links / noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. the page instead). - Decided who should rank — default is your original; any exception is written down as deliberate.
- Chose the instruction:
noindex(Google), canonical (Bing / hint), ornoindex+ canonical (both). - Put the tag requirement in the syndication agreement — exact tag, scope, present at publish.
- For Google Search: partner uses
<meta name="robots" content="noindex">(orgooglebot). - For Google News only: partner uses
<meta name="googlebot-news" content="noindex">. - For a preserved link back: implemented as
noindex,follow. - For Bing: partner uses a cross-domain
rel=canonicalback to your original; excerpt-plus-link where possible. - Published and indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. your original first (or simultaneously).
- Verified the tag is live in the reprint’s rendered
<head>on launch day, and that the reprint’s URL isn’t blocked in the partner’srobots.txt(noindexonly works if Google can crawl the page and see the tag). - Scheduled a 2–4 week recheck: headline SERP check + Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. Google-selected canonicalA Google Search Console Page Indexing status: you declared a canonical for this URL, but Google overrode your choice, picked a different page as the canonical, and indexed that one instead..
- Documented the deal (partner, URLs, chosen instruction, date) for the next audit.
Ready-to-copy AI prompts
Prompts for working through a syndication decision. Paste your specifics in the
[brackets].
Pick the right instruction
I run [your site] and I’m syndicating an article to [partner]. My original is at [URL]. I want my original to be the version that ranks. Walk me through whether I should ask the partner for a cross-domain rel=canonicalA rel=\"canonical\" annotation — in the HTML <head> or an HTTP Link header — that tells search engines which URL is the preferred version of duplicate or near-duplicate content., a noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed., or both — based on Google’s current guidance (which recommends noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. for syndication since ~May 2023) and Bing’s (which still prefers canonical). Explain the tradeoff of each and give me the exact meta tag to request.
Draft the contract clause
Draft a short syndication-agreement clause requiring the partner to add [noindex / a canonical pointing to my original at [URL]] on their reprint, present at publish time. Specify the exact tag and make clear whether it should scope to Google News only or all of search. Keep it plain-English and enforceable.
Diagnose a partner outranking me
My original article at [URL] is being outranked by a syndication partner’s copy at [partner URL] for the query [query]. Give me a step-by-step diagnostic: what to check in Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. (URL InspectionA Google Search Console feature that reports how Google sees one specific URL on a property you own. By default it shows the last-indexed snapshot; a separate \"Test live URL\" mode fetches the current version., Google-selected canonicalA Google Search Console Page Indexing status: you declared a canonical for this URL, but Google overrode your choice, picked a different page as the canonical, and indexed that one instead.), how to confirm whether the partner’s page has a noindex or canonical, and the exact fix to request. Note that noindex needs a recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial. before it takes effect.
Explain the mechanism to a stakeholder
Explain, for a non-technical publisher, why Google now recommends noindex over a canonical tagA rel=\"canonical\" annotation — in the HTML <head> or an HTTP Link header — that tells search engines which URL is the preferred version of duplicate or near-duplicate content. for syndicated contentSyndicated content is the deliberate, usually licensed republishing of the same article on other sites — wire pickups, partner reprints, or cross-posting. Because both parties know it's a copy, the SEO job is deciding which URL search engines should show, not eliminating a duplicate. — specifically the “pages are often very different” reasoning and how that makes a cross-domain canonical unreliable. Contrast it with ordinary duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling. so they understand syndication is a policy decision, not a bug.
Mental models for syndicated content
1. A policy decision, not a duplicate-content bug. The reprint exists intentionally. The first question is who should be eligible to rank, not which accidental URL variant needs cleanup. Put the indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. requirement into the syndication agreement.
2. Three paths, one explicit choice. A partner can point a cross-domain canonical to the original, noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. the reprint, or leave both indexable. “Do nothing” is still a policy choice—it accepts the risk that the partner becomes the selected URL.
3. Hint versus directive. A canonical asks an engine to treat pages as duplicates and
consolidate on the original. Because partner chrome, ads, navigation, and edits can make
the pages materially different, Google may ignore that hint. noindex removes the
reprint from eligibility after recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial., which is why Google now prefers it.
4. Scope the directive to the goal. googlebot-news: noindex targets Google News;
googlebot: noindex or robots: noindex covers regular Search too. Choose scope based on
where cannibalization matters instead of copying a tag without understanding it.
5. Engine guidance can legitimately diverge. Google prefers partner noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed.; Bing
still recommends a canonical where the agreement allows and suggests excerpt syndication.
When both matter, noindex plus a canonical can express both preferences.
6. Publish first reduces ambiguity, not risk to zero. Getting the original crawled and indexed before partners publish gives engines an origin signal, but it does not override a noncompliant partner or guarantee canonical selectionHow search engines pick one canonical URL among duplicates and consolidate signals onto it..
7. Syndication for links is a different problem. Replicating articles widely mainly to place keyword-rich backlinks is link-scheme territory. Canonical/noindex mechanics do not turn a manipulative campaign into legitimate editorial syndication.
Syndication control reference
Choose the path
| Partner setup | Google effect | Bing effect | Use when | Main risk |
|---|---|---|---|---|
| Cross-domain canonical to original | A hint Google may ignore when pages differ; no longer Google’s preferred syndication fix | Bing’s stated preference where agreements allow | Bing matters and the partner will supply a faithful canonical | Partner can remain indexable and outrank the original |
noindex on reprint | Google’s preferred solution; removes the partner copy after recrawlCrawl frequency is how often a search engine comes back to re-fetch a page it already knows about. Popular pages that change often get refreshed many times a day; stable pages can go weeks or months between crawls — and you influence it indirectly, not by setting a dial. | Removes the copy where Bing honors the directive too | The original must be the only version eligible to rank | Referral visibility from search is intentionally lost |
noindex + canonical | Directive removes the reprint while the canonical still identifies the original | Preserves Bing’s canonical hint | Both engines matter and the agreement supports both tags | Requires the partner to maintain both signals correctly |
| Nothing | Google selects a canonical without your instruction | Bing also resolves the ambiguity itself | You genuinely accept the partner ranking | A larger partner may win your story |
| Excerpt + source link | Keeps less duplicate text in circulation | Explicitly encouraged by Bing where possible | The partnership can deliver reach without a full reprint | The excerpt still needs a clear agreement and attribution |
Choose the noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. scope
| Goal | Tag on the partner copy | Result |
|---|---|---|
| Keep reprint out of Google News only | <meta name="googlebot-news" content="noindex"> | Partner can remain in regular Search but not Google News |
| Keep reprint out of Google Search and News | <meta name="googlebot" content="noindex"> | Google-specific exclusion across those surfaces |
| Keep reprint out of search engines that honor robots meta | <meta name="robots" content="noindex"> | Broad crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. directive |
| Preserve a followable attribution link while excluding the page | noindex,follow | Page is removed after recrawl; links are not explicitly nofollowed, though recrawl can decline over time |
Precondition: any of the above only works if Google can crawl the page and see
the tag — the reprint’s URL must not be blocked in the partner’s robots.txt, and
must otherwise be reachable to the crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.. A blocked or gated page means the
noindex is never seen.
Agreement clauses to settle before publication
- Which URL is intended to rank.
- Which engine/surface scope the indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. directive covers.
- Whether a canonical is also required for Bing.
- Whether the partner publishes a full article or excerpt.
- How quickly the partner must fix a missing tag.
- Whether the original publishes and indexes before the reprint goes live.
Test yourself: Syndicated Content and SEO
Five quick questions on how syndication and canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it. interact. Pick an answer for each, then check.
Resources worth your time
My related writing
- Google Uses ~40 Canonicalization Signals — my deep dive on how Google actually selects a canonical, and the “it’s a hint, not a rule / strong signal but overridable” framing that underpins why canonical is the least reliable tool for syndication.
- Canonical Tags Explained: Why They Matter For SEO — the Ahrefs implementation guide (byline Joshua Hardwick) I review: golden rules, common mistakes, and how to test a canonical.
- The Beginner’s Guide to Technical SEO — where canonicalizationHow search engines pick one canonical URL among duplicates and consolidate signals onto it. and indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. control sit in the bigger picture.
Official
- Fix canonicalization issues — syndicated content section (Google) — the current, load-bearing “noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed., not canonical” guidance.
- Avoid article duplication in Google News (Google) — the exact
Googlebot-News/GooglebotnoindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. tags. - Does Duplicate Content Hurt SEO and AI Search Visibility? (Bing, Dec 2025) — Bing’s canonical-preferred stance and AI-search angle.
From around the industry
- Why Noindexing Syndicated Content Is The Way — a 3,000-URL case study (Glenn Gabe, GSQI) — the independent tracking study behind “no rhyme or reason” in which copy Google surfaces.
- Google no longer recommends canonical tags for syndicated content (Search Engine Land) — coverage of the ~May 2023 doc change.
- Google Recommends Noindex For Syndicated News Content (Search Engine Journal) — Sullivan’s “focus on your goal, not the mechanism” framing.
- Google explains why syndicators may outrank original publishers (Search Engine Land) — Sullivan on the “difficult to identify the originating source” problem.
- Google On When To Use Noindex & Canonical Tags (Search Engine Roundtable) — the dated April 22, 2023 Mueller post naming syndication as a noindex case.
- Google warns against misusing links in syndication & large-scale article campaigns (Search Engine Land) — where syndication crosses into a link scheme.
- What is Article Syndication (Ahrefs Glossary) — a plain-language definition of the term.
Syndicated Content
Syndicated content is the deliberate, usually licensed republishing of the same article on other sites — wire pickups, partner reprints, or cross-posting. Because both parties know it's a copy, the SEO job is deciding which URL search engines should show, not eliminating a duplicate.
Related: Canonical Tag, Duplicate Content, Noindex
Syndicated Content
Content syndication is when the same article runs on more than one website on purpose — a wire service like the AP or Reuters feeding its members, a publisher licensing stories to a partner like Yahoo or MSN, a press release distributed at scale, or a company cross-posting its own blog to Medium or LinkedIn. The defining trait is that it’s intentional and usually contractual: both the original publisher and the partner know the content is a copy.
That’s what separates it from ordinary duplicate contentThe same or very similar primary content reachable at more than one URL. There's no general duplicate content penalty — the real costs are possible signal dilution, the wrong URL getting chosen, and less-efficient crawling., which is an accident of URL parametersThe `?key=value` data tacked onto the end of a URL after a question mark — used for tracking, sessions, filtering, sorting, and search — and one of the biggest sources of duplicate URLs and wasted crawling in SEO., staging environments, or scraping. The fix is different too. With accidental duplication you eliminate the extra copy; with syndication both copies are supposed to exist, so the job is to tell search engines which one should appear in results.
That decision is one of three: a cross-domain rel=canonical on the reprint pointing back to the original, a noindex on the reprint, or doing nothing and letting Google pick. As of a May 2023 guidance change, Google explicitly recommends noindexNoindex is a directive that tells search engines to keep a page out of their index, so it won't appear in search results. It works only on pages a crawler can actually fetch — a page blocked in robots.txt can never be noindexed. over canonical for syndication, because syndicated pages are “often very different” once template chrome, ads, and navigation differ — which makes the canonical unreliable. Get the decision wrong (or skip it) and a bigger syndication partner can end up outranking you for your own story.
Related: Canonical Tag, Duplicate Content, Noindex
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.
Search Console
sampleGA4 traffic (28d)
sampleCloudflare traffic (7d)
sampledCrUX field data (28d, phone)
sampleGoogle NLP entities
localChangelog
Updated Jul 18, 2026.
Editorial summary and recorded change details.Summary
Reviewed against a fresh research packet and re-verified Google's current syndication guidance live — the article's noindex-over-canonical position holds. Added the crawlability precondition for noindex: it only works if Google can crawl the reprint and see the tag, so a partner's robots.txt block or gated page defeats it.
Change details
-
Added Google's documented precondition that noindex only takes effect if the page isn't blocked by robots.txt and is otherwise reachable to the crawler, with a matching check added to the playbook runbook, the SOP, the checklist, the cheat sheet, and the decision tree's noindex-everything result.
Full comparison unavailable — no prior snapshot was archived for this revision.