Paywalls i SEO
How to zachować paywalled i registration-gated treść indeksowalny bez cloaking — flexible sampling, isAccessibleForFree/cssSelector znaczniki, the JavaScript-paywall trap, i metering strategy.
Języki
A paywall doesn't inherently hurt SEO — Google ma no bias wobec gated treść, i the biggest paywalled publishers rank fine. co hurts jest Google nie będąc able to see enough treść to zrozum strona. The supported fix jest flexible sampling: let Googlebot crawl the pełny artykuł, then declare the gated part z dane strukturalne (isAccessibleForFree plus a cssSelector). że's an explicit, sanctioned exception to cloaking — cloaking jest o intent to deceive; ten jest a declared mechanism. używać metering (start around 6–10 free artykuły/month) lub lead-in, gate serwer-side (nie z JavaScript że just hides treść in the DOM), give login strony unique copy, i nigdy używać robots.txt to hide private URLs.
TL;DR — A paywall (subscription, one-time payment, lub just a registration/login gate) doesn’t automatically hurt twój SEO. Google ma a supported way to handle it called flexible sampling: you let Googlebot read the whole artykuł, then używać a bit of dane strukturalne to tell Google który part jest gated. Done że way, showing Google the pełny artykuł podczas gdy czytelnicy see a truncated version jest nie cloaking — it’s an approved exception.
robić paywalls hurt SEO?
Google obsługuje paywalled treść gdy crawlers może access it i the implementacja używa the udokumentowany paywall structured-data pattern. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Paywalled content structured data Google’s flexible-sampling guidance describes metering i lead-in approaches, nie a ranking guarantee. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Flexible sampling
nie on ich own. ten jest the pierwszy thing to get straight, ponieważ half the poradniki out there frame paywalls as an SEO problem to być minimized. They aren’t. Google ma no bias wobec paywalled treść — the New York Times, the Wall Street Journal, the Financial Times, i the Washington Post wszystkie sit behind paywalls i rank prominently dla exactly the stories they gate.
co robi hurt rankings jest Google nie będąc able to see enough of twój treść to understand co the strona jest o. If a bot tylko ever sees a two-zdanie teaser, it może tylko rank you dla tamte two zdania. So the whole game z paywalls i SEO jest ten: let the wyszukiwarka przeczytaj pełny artykuł, podczas gdy normal visitors nadal hit the gate.
One thing ten znaczniki jest nie: a promise. Getting isAccessibleForFree i the
rest of the znaczniki exactly right doesn’t guarantee indeksowanie, ranking, lub a rich
wynik — Google’s own structured-data documentation says plainly it doesn’t
guarantee dowolny funkcja będzie pokazywać up in wyniki wyszukiwania. co the znaczniki robi jest
usuń cloaking risk of showing crawlers więcej niż użytkownicy see; it doesn’t
manufacture rankings by itself.
The supported way to robić it: flexible sampling
Google’s model jest called flexible sampling, i it ma two flavors:
- Metering — visitors get a quota of free artykuły (Google suggests starting around 6–10 per month) przed the paywall kicks in.
- Lead-in — you pokazywać the opening of an artykuł, then gate the rest.
On top of whichever you choose, you dodawać a mały piece of dane strukturalne to the strona że tells Google, “this section is behind a paywall.” że’s the label że makes everything legitimate.
Isn’t showing Google the pełny artykuł cheating?
ten jest the question everyone asks, i the answer jest no — ponieważ you declared it. Cloaking (the bad thing) jest gdy you pokazywać wyszukiwarki różny treść than użytkownicy aby deceive them i manipulate rankings. Flexible sampling jest the opposite: you’re openly telling Google, przez dane strukturalne, “hey, rzeczywisty użytkownicy see something więcej limited than co you’re crawling.” Google’s own spam polityka carves paywalls out of the cloaking definition by nazwa, as long as you follow the flexible-sampling guidance i let Google see the pełny treść.
The one mistake to avoid
Don’t build twój paywall by wysyłka the entire artykuł in the strona’s HTML i just hiding it z JavaScript lub CSS until someone logs in. It feels easier, ale it backfires: anyone może turn off JavaScript i read twój paid treść dla free, screen czytelnicy będzie przeczytaj “hidden” tekst aloud, i Google może’t reliably tell który part you meant to gate. The right way jest to gate it on the serwer — tylko wysyłać the pełny artykuł once you’ve confirmed the person jest logged in lub subscribed.
Want the pełny mechanics — the dokładny dane strukturalne, the metering liczby, the JavaScript trap, i why login strony cause ich own problems? Switch to the Advanced tab.
TL;DR — Paywalls don’t inherently hurt rankings; Google będąc unable to see twój treść robi. The supported model jest flexible sampling — metering lub lead-in — declared z dane strukturalne (
isAccessibleForFree: falseplus ahasPart/cssSelectormarking the gated sekcja, class selectors tylko). że declaration jest co makes serving Googlebot the pełny artykuł nie cloaking: cloaking wymaga intent to manipulate i mislead, i Google’s spam polityka explicitly carves paywalls out of że definition. Gate serwer-side (the 2025 doc update i Mueller’s screen-czytelnik caution oba target the same JS-hiding mistake), give login strony unique copy, nigdyrobots.txtprivate URLs, i używaćnoarchiveto stop a cached copy leaking the pełny tekst. Registration walls używać the same znaczniki as paid ones.
co actually causes ranking problems (it isn’t the gate)
Paywall eligibility depends on crawlable treść i accurate znaczniki; the presence of a paywall alone jest nie udokumentowany as a penalty. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Paywalled content structured data Sampling choices pozostawać publisher decisions z użytkownik i firma tradeoffs. Evidence for this claim Official or primary documentation supporting the adjacent article claim, with scope limited to the source's published description. Scope: No ranking guarantee or undisclosed system mechanics are inferred beyond the cited source. Confidence: high · Verified: Google: Flexible sampling
Google ma no penalty dla paywalled treść, i ten artykuł’s parent hub says as much: gated treść jest fine as long as Google może read it przez the supported approach. The awaria mode jest upstream of ranking — it’s comprehension. If Googlebot tylko ever sees a teaser, że teaser jest wszystkie it może index i rank you dla. każdy technique below exists to solve one problem: let the engine przeczytaj whole thing, podczas gdy unauthenticated humans nadal hit the gate.
Two granice worth stating plainly, since it’s łatwy to overreach in either
direction. pierwszy, ten znaczniki jest a narzędzie dla treść you want zindeksowany poniżej a
declared gate — nie a mechanism dla exposing treść you don’t want zindeksowany at
wszystkie. Genuinely private account/admin URLs są a różny case (see the decision
tree below): tamte get noindex lub an authentication redirect, nie
isAccessibleForFree. Second, prawidłowy znaczniki i pełny crawl access są nie a
ranking guarantee. Google’s structured-data guidelines say directly że “Google
robi nie guarantee że funkcje że consume dane strukturalne będzie pokazywać up in
wyniki wyszukiwania” — the znaczniki jest the declaration że zachowuje you out of the
cloaking bucket, nie a promise of indeksowanie, ranking, ruch, lub a wynik z elementami rozszerzonymi.
Historically ten jest gdzie the biggest cautionary tale comes z. gdy the Wall Street Journal pulled out of Google’s old pierwszy Click Free program in 2017, it reported a ~44% drop in Google ruch z wyszukiwania — nie ponieważ paywalls są penalized, ale ponieważ Google mógł no longer see the artykuły at wszystkie. (więcej on pierwszy Click Free below; it jest history, nie current polityka.)
Flexible sampling: metering i lead-in
The current, active model jest flexible sampling, laid out in Google’s Flexible Sampling guidelines. Google describes two sampling types: “metering, który zapewnia użytkownicy z a quota of artykuły to consume przed requiring użytkownicy to subscribe lub log in, po który paywalls będzie start appearing; i lead-in, który oferty a portion of an artykuł’s treść bez it będąc shown in pełny.”
The liczby że matter, wszystkie z Google’s own doc:
- Prefer monthly ponad daily metering. Google: “In general, we think że monthly, zamiast daily metering zapewnia więcej flexibility i a safer environment dla testing.” A one-unit change jest far mniej jarring at 10 monthly samples than at 3 daily ones.
- Start around 6–10 free artykuły/month. “As a starting point dla twój explorations, we encourage you to zapewniać 10 artykuły per month… dla najbardziej daily news publishers, we expect the wartość to fall między 6 i 10 artykuły per użytkownik per month.”
- Watch the exposure ceiling. “nasz analiza pokazuje że general użytkownik satisfaction starts to degrade significantly gdy paywalls są shown więcej niż 10% of the time (który generally means że o 3% of the odbiorcy ma był exposed to the paywall).”
- Lead-in jest a good practice. Showing the pierwszy kilka zdania above the paywall lets użytkownicy “experience the value of the content.”
None of te liczby jest a mandate. Google says directly: “There jest no single wartość dla optimal sampling w całym różny firmy” — the 6–10/month figure jest a starting point Google gives specifically dla daily news publishers, i even że comes z “we leave the dokładny liczba to the discretion of individual publishers, who są best positioned to zrozum particular demands of ich firmy.” Treat it as a tested starting range, nie a reguła to copy verbatim.
The poniżej-appreciated point: metering jest nie purely a monetization dial. Google opens the doc noting że “even minor changes to the current sampling levels mógł degrade doświadczenie użytkownika i, as użytkownik access jest restricted, unintentionally impact artykuł ranking in Google Search.” Tightening the meter może quietly cost you rankings.
Why ten isn’t cloaking — the reasoning, nie just the reguła
ten jest the load-bearing part of the whole topic, i najbardziej poradniki assert the conclusion (“paywalls aren’t cloaking if you use structured data”) bez showing why. Here’s the rzeczywisty reasoning, straight z Google’s spam polityki.
zacznij od the definition. Cloaking jest “the practice of presenting różny treść to użytkownicy i wyszukiwarki z the intent to manipulate search rankings i mislead użytkownicy.” The load jest on że clause — intent to manipulate i mislead. A paywall isn’t trying to trick anyone; it’s monetizing treść, i it’s declaring the difference in treatment przez znaczniki.
Then the explicit carve-out, in the same polityka: “If you operate a paywall lub a treść-gating mechanism, we don’t consider ten to być cloaking if Google może see the pełny treść of co’s behind the paywall just like dowolny person who ma access to the gated material i if you follow nasz Flexible Sampling general guidance.”
So the exception ma two conditions: (1) Google sees the same pełny treść a paid subscriber by, i (2) you follow flexible sampling — który w praktyce means the dane strukturalne below. Google’s flexible-sampling doc reinforces the same logic: “Enclose paywalled treść z dane strukturalne aby pomagać Google differentiate paywalled treść z the practice of cloaking, gdzie the treść served to Googlebot jest różny z the treść served to użytkownicy.” The structured data jest the declaration że turns “different content for bots” z deception do a disclosed, sanctioned mechanism.
Implementing the dane strukturalne
The znaczniki lives in Google’s Subscription i paywalled treść doc. Two właściwości robić the działać:
isAccessibleForFree(Boolean, required) — whether the treść jest free lub gated. Google’s own właściwość reference marks ten the required one; ustawić it on the top-levelCreativeWork/NewsArticlenode i on każdy gated sekcja.hasPart(recommended, nie required) — an tablica ofWebPageElementobiekty, one per gated sekcja, każdy z jego ownisAccessibleForFree: falsei acssSelectorpointing at the class you wrapped the gated HTML in. ten jest how you tell Google który part of the piece jest gated gdy it’s a sekcja rather than the whole thing; it’s the recommended way to get sekcja-level precision, nie a second required właściwość alongside the top-level flag.
A minimal NewsArticle looks like ten:
{
"@context": "https://schema.org",
"@type": "NewsArticle",
"isAccessibleForFree": false,
"hasPart": {
"@type": "WebPageElement",
"isAccessibleForFree": false,
"cssSelector": ".paywall"
}
}Three implementacja details people trip on:
- Class selectors tylko. The
cssSelector“references the class nazwa że you ustawić in the HTML.” używać.paywall— nie an ID (#paywall), nie a descendant lub atrybut selector. - Multiple gated sekcje używać an tablica of
hasPartobiekty, każdy z jego own class-oparty selector. Don’t nest the gated sekcje inside każdy other. - It’s nie just dla news. The znaczniki jest supported on dowolny
CreativeWorksubtype —Article,NewsArticle,Blog,Comment,Course,HowTo,Message,Review,WebPage. The broader dane strukturalne guidance treatsisAccessibleForFreeas a generalCreativeWorkwłaściwość, nie a news-tylko one. - poprawny znaczniki doesn’t guarantee a wynik. Even fully prawidłowy, correctly-nested znaczniki tylko makes Google eligible to understand twój gating — it isn’t a ranking lub rich-wynik guarantee. Treat the znaczniki as the mechanism że zachowuje you out of the cloaking bucket, nie a promise of dowolny specific outcome.
Registration walls użyj identical znaczniki. Google doesn’t distinguish “pay to access” from “register to access” at the schemat level. John Mueller said as much on Search Off the Record: the mechanism “mógł być maybe you wymagać a login, maybe you wymagać a payment, maybe po a certain liczba of iterations you’re like, ‘Oh, ten jest enough free treść.’ Now you mieć to pay dla it… It może just być something like a login lub niektóre other mechanism że basically limits the visibility of the treść.” If you gate it, mark it — paid lub nie. He even flags A/B pricing tests as a prawidłowy powód: “if you mieć something like różny thresholds gdzie you say niektóre people get to view five strony dla free i others mieć the whole treść available dla free ponieważ you’re doing A/B testing… then you’d want to używać a paywall structured data.”
The JavaScript-paywall trap
Here’s the single najbardziej common rzeczywisty-world mistake, i it’s distinct z “forgetting the dane strukturalne.” A lot of paywall solutions ship the pełny artykuł in the HTML the serwer wysyła, then używać JavaScript to hide it until subscription status jest confirmed. Google explicitly warned wobec ten in a 2025 addition to jego JavaScript troubleshooting doc: “niektóre JavaScript paywall solutions obejmować the pełny treść in the serwer odpowiedź, then używać JavaScript to hide it until subscription status jest confirmed. ten isn’t a niezawodny way to limit access to the treść. upewnij się twój paywall tylko zapewnia the pełny treść once the subscription status jest confirmed.”
Why it’s bad on three fronts:
- It’s trivially bypassable. Disable JavaScript i the “hidden” artykuł jest right there in the źródło. You’re nie actually gating anything.
- It muddies the cloaking exception. If the pełny tekst jest sitting in the DOM dla everyone, Google może’t cleanly tell który treść był meant to być gated — który jest the whole thing the structured-data declaration jest supposed to make jasny.
- It’s an accessibility problem. Mueller raised exactly ten on Search Off the Record: “gdy a użytkownik looks at twój strona, you don’t load the treść do the HTML, ale rather you upewnij się że it’s really nie załadowany do the strona’s DOM so że, if a przeglądarka ma something like… a screen czytelnik, że the screen czytelnik doesn’t go off i read wszystkie of ten tekst że you’re trying to hide… upewnij się you don’t load it do the przeglądarka i używać JavaScript to turn it on, ale rather że it’s really tylko served to the użytkownik gdy you want to make it available.” The 2025 doc update i Mueller’s caution są the same mistake seen z two angles.
The fix jest serwer-side gating: confirm subscription/login status on the serwer,
i tylko obejmować the pełny artykuł in the odpowiedź dla authenticated użytkownicy. Then
warstwa isAccessibleForFree/cssSelector on top so Googlebot — który jest allowed
to see the pełny tekst poniżej flexible sampling — nadal gets everything, podczas gdy
unauthenticated humans genuinely don’t. ten jest również gdzie paywalls intersect z
mobilny-pierwszy indeksowanie: Google crawls i evaluates the mobilny version, so the pełny
gated treść ma to być present in the mobilny serwer odpowiedź too, nie just komputer stacjonarny.
Login strony i registration gates: the quieter pitfalls
Two distinct problems pokazywać up around login/registration, oba z the same Search Off the Record episode.
Generic login strony get folded do duplicates. Mueller: “if you mieć a bardzo generic login strona, we będzie see wszystkie of te URLs że pokazywać że login strona, że redirect to że login strona, as będąc duplicates… We’ll fold them together as duplicates, i we’ll focus on indeksowanie the login strona… If someone jest searching dla twój service… the tylko thing… they find in search jest like, ‘Here’s how to log in,’ że może być a kind of a weird experience dla them.” The fix jest to give login strony unique contextual copy per service, so they’re nie wszystkie identical.
Don’t robots.txt private URLs. ten one contradicts a common intuition.
Mueller: “whether wszystkie of ten powinien just być blocked by robots.txt, który jest another
common strategy… The problem, I think, z doing że jest the URLs mógł become
indeksowalny so we wouldn’t see the treści of the login strona… if it’s private
treść, serve it z a noindex lub redirect it to a login strona somewhere. Don’t
używać robots.txt.” A robots-blocked URL może nadal być zindeksowany as a bare, contentless
URL — często worse than a clean noindex. (ten jest genuinely-private treść, który
jest a różny case z paywalled-ale-powinien-być-zindeksowany; don’t confuse the two.)
Testing i the “leaky” worry
Test z the wyniki z elementami rozszerzonymi Test. Google
added paywalled-treść obsługiwać
to the wyniki z elementami rozszerzonymi Test in October
2023, so it validates isAccessibleForFree/cssSelector on a live URL, testing as
Googlebot komputer stacjonarny lub smartphone. As Mueller put it back in a 2020 office-hours,
“you by użyj wyniki z elementami rozszerzonymi test, like dowolny other kind of dane strukturalne… the
tricky part z niektóre of te paywall implementacje jest że Googlebot, of course,
needs to być able to see the pełny treść.”
The self-audit trick: otwarty an incognito window (logged out of everything), search dla twój own marka lub service, i see co pokazuje. Mueller’s advice — “If the top wynik jest something like a login strona i there’s no information on ten strona at wszystkie otherwise, then probably że’s something że you może poprawić.”
jest showing Googlebot the pełny artykuł “leaky”? No. Danny Sullivan, Google’s
Search Liaison, addressed the recurring worry że ten exposes paid treść:
“nasz system jest looking to być shown the pełny treść, if a publisher wants to robić
że. If they robić, we understand więcej o it. If we understand więcej, then we może
być able to pokazywać it dla więcej zapytania gdzie it’s relevant,” and “Since tylko we są
seeing ten, there’s nothing ‘leaky’ as you są suggesting.” The rzeczywisty leak vector,
he noted, jest the cached copy — solved z noarchive, a oddzielny control z
the paywall znaczniki itself.
Sullivan’s remarks są relayed via wyszukiwarka Roundtable’s coverage;
treat them as reported zamiast a pierwszy-party transcript.
Bing’s approach
Bing’s subscription i paywall guidance
(Fabrice Canel, może 2022) jest structurally similar ale nie schemat-centric. jego
three points: (1) let Bingbot crawl the pełny gated treść, (2) używać
noarchive/nocache (lub the X-Robots-Tag: noarchive header) so cached copies
don’t leak, i (3) verify the crawler jest genuinely Bingbot by checking the
requesting IP wobec Bing’s opublikowany ranges — nie by trusting the użytkownik-agent
ciąg znaków, który anyone może spoof. There’s no opublikowany Bing equivalent to
isAccessibleForFree/cssSelector; Bing’s model jest crawl-access-plus-pamięć podręczna-control,
gdzie Google’s jest znaczniki-centric. Don’t assume funkcja parity.
pierwszy Click Free — history, nie polityka
You’ll nadal see blog posts i forum answers describing pierwszy Click Free as if it’s current. It isn’t. Google retired it in October 2017, zastępowanie it z flexible sampling. Richard Gingras, then Google’s VP of News: “pierwszy, Flexible Sampling będzie zastępować pierwszy Click Free. Publishers są in the best position to determine co level of free sampling działa best dla them.” pierwszy Click Free miał required participating publishers to let Google-referred visitors read a ustawić liczba of artykuły a day (commonly three) even past ich own paywall. Flexible sampling handed że decision back to publishers. If you see FCF cited as something you może opt do today, że guidance jest eight-plus years stale.
gdzie ten sits in news SEO
Paywall handling jest one piece of the broader News & odkrywać SEO picture — alongside news sitemaps, Google News/Top Stories eligibility, odkrywać, i syndication (canonical vs. noindex). If you’re a publisher, get twój paywall znaczniki i twój syndication polityka sorted przed either one quietly costs you indexation lub attribution.
AI summary
A condensed take on the Advanced version:
- Paywalls don’t inherently hurt SEO. Google ma no bias wobec gated treść; the biggest paywalled publishers rank fine. co hurts jest Google będąc unable to see enough treść to zrozum strona — i the znaczniki itself jest a declaration, nie a ranking guarantee (Google’s own docs say dane strukturalne doesn’t guarantee dowolny funkcja będzie pokazywać up in wyniki wyszukiwania).
- Flexible sampling jest the supported model (nie the retired pierwszy Click Free, gone since Oct 2017): metering (start ~6–10 free artykuły/month, prefer monthly ponad daily) lub lead-in (pokazywać the opening, gate the rest). Google jest explicit there’s “no single wartość dla optimal sampling w całym różny firmy” — 6–10/month jest a daily-news starting point, nie a universal reguła. użytkownik satisfaction degrades past ~10% paywall-exposure; tightening the meter może even cost rankings.
- dane strukturalne jest the mechanism:
isAccessibleForFree: false(the required właściwość) on the artykuł node, plus a recommendedhasPart/WebPageElementz a class-opartycssSelectordla sekcja-level precision. działa on dowolnyCreativeWorksubtype, nie just news. It’s dla treść you want zindeksowany poniżej a declared gate — genuinely private URLs getnoindexinstead, nie ten znaczniki. - Why it isn’t cloaking: cloaking wymaga intent to manipulate i mislead; Google’s spam polityka explicitly carves out paywalls gdy Google sees the pełny treść i you follow flexible-sampling guidance. The znaczniki jest the declaration.
- Registration/login walls użyj identical znaczniki as paid paywalls — Google doesn’t distinguish pay-vs-register at the schemat level (per Mueller).
- The JS-paywall trap: don’t ship the pełny artykuł in the HTML i hide it z JS/CSS — it’s bypassable, muddies the cloaking exception, i screen czytelnicy read the “hidden” tekst. Gate serwer-side; the pełny treść musi być in the mobilny odpowiedź too (mobilny-pierwszy indeksowanie).
- Login-strona pitfalls: generic login strony get folded as duplicates (give them
unique copy); nigdy
robots.txtprivate URLs (używaćnoindex/redirect). - pamięć podręczna leak jest a oddzielny control:
noarchive/nocachestops a cached copy exposing gated tekst (per Danny Sullivan). Bing’s model jest crawl-access + pamięć podręczna control + IP verification, z noisAccessibleForFreeequivalent. - Test z the wyniki z elementami rozszerzonymi Test (paywall obsługiwać since Oct 2023) i Mueller’s incognito self-audit.
Official documentation
Primary-źródło guidance z the wyszukiwarki.
- Flexible Sampling — the core model: metering vs. lead-in, the 6–10 artykuły/month starting point, the 10%-exposure ceiling, i the cloaking-differentiation rationale.
- Subscription i paywalled treść znaczniki —
isAccessibleForFree,hasPart/WebPageElement, i the class-opartycssSelector. - Spam polityki — Cloaking — the cloaking definition i the explicit paywall carve-out.
- Fix search-powiązany JavaScript problems — the 2025 JavaScript-paywall guidance.
- Google common crawlers lista — the rzeczywisty crawler użytkownik-agents (używany below to debunk the fabricated “Googlebot Subscriber” twierdzenie).
- Driving the future of digital subscriptions — the 2017 pierwszy Click Free → Flexible Sampling transition.
- wyniki z elementami rozszerzonymi Test — validates paywall dane strukturalne on a live URL.
Bing / Microsoft
- SEO dobra praktyka dla subscription-oparty i paywall treść — Fabrice Canel, może 2022: crawl access, pamięć podręczna control, i IP-oparty Bingbot verification.
cytaty z the źródło
On-the-record statements z Google i Bing. każdy link deep-links to the quoted passage gdzie the źródło strona obsługuje it.
Google — why paywalls aren’t cloaking (the load-bearing cytaty)
- “Cloaking refers to the practice of presenting different content to users and search engines with the intent to manipulate search rankings and mislead users.” — Google spam polityki. Jump to cytat
- “If you operate a paywall or a content-gating mechanism, we don’t consider this to be cloaking if Google can see the full content of what’s behind the paywall just like any person who has access to the gated material and if you follow our Flexible Sampling general guidance.” Jump to cytat
- “Enclose paywalled content with structured data in order to help Google differentiate paywalled content from the practice of cloaking, where the content served to Googlebot is different from the content served to users.” Jump to cytat
Google — flexible sampling i metering
- “There are two types of sampling we advise: metering, which provides users with a quota of articles to consume before requiring users to subscribe or log in, after which paywalls will start appearing; and lead-in, which offers a portion of an article’s content without it being shown in full.” Jump to cytat
- “In general, we think that monthly, rather than daily metering provides more flexibility and a safer environment for testing.” Jump to cytat
- “As a starting point for your explorations, we encourage you to provide 10 articles per month to Google search users and iterate from there… for most daily news publishers, we expect the value to fall between 6 and 10 articles per user per month.” Jump to cytat
- “Our analysis shows that general user satisfaction starts to degrade significantly when paywalls are shown more than 10% of the time (which generally means that about 3% of the audience has been exposed to the paywall).” Jump to cytat
Google — the JavaScript-paywall trap
- “Some JavaScript paywall solutions include the full content in the server response, then use JavaScript to hide it until subscription status is confirmed. This isn’t a reliable way to limit access to the content. Make sure your paywall only provides the full content once the subscription status is confirmed.” Jump to cytat
Richard Gingras, VP of News, Google (Oct 2017)
- “First, Flexible Sampling will replace First Click Free. Publishers are in the best position to determine what level of free sampling works best for them.” przeczytaj announcement
John Mueller, Google — Search Off the Record (Sep 2025)
- On registration vs. payment gates: “It also doesn’t have to be something that’s behind a clear payment thing. It can just be something like a login or some other mechanism that basically limits the visibility of the content.”
- On the DOM/screen-czytelnik caution: “you make sure that it’s really not loaded into the page’s DOM so that, if a browser has something like… a screen reader, that the screen reader doesn’t go off and read all of this text that you’re trying to hide.”
- On private URLs: “if it’s private content, serve it with a noindex or redirect it to a login page somewhere. Don’t use robots.txt.” pełny transcript (PDF)
John Mueller, Google — SEO office-hours (Dec 2020)
- “Essentially you would use the rich results test, like any other kind of structured data. I think the tricky part with some of these paywall implementations is that Googlebot, of course, needs to be able to see the full content so that we can understand what it is that we should be showing your site for.” Coverage (wyszukiwarka Journal)
Danny Sullivan, Google Search Liaison — the “not leaky” clarification
- “Our system is looking to be shown the full content, if a publisher wants to do that. If they do, we understand more about it. If we understand more, then we might be able to show it for more queries where it’s relevant.” i “Since only we are seeing this, there’s nothing ‘leaky’ as you are suggesting.” Coverage (wyszukiwarka Roundtable)
który paywall setup robić I need?
Paywall implementacje differ mostly on how you gate i co you want zindeksowany. Walk przez it — the leaf tells you który znaczniki (if dowolny) i który control applies.
Choosing the right gating + markup approach
co nie to robić z paywalls
1. Treating pierwszy Click Free as current polityka. Plenty of stale posts describe pierwszy Click Free as if you może nadal opt in. Google retired it in October 2017 i replaced it z flexible sampling. Fix: design around metering/lead-in i dane strukturalne; if you see FCF cited as live guidance, ignore it.
2. Hiding the pełny artykuł z JavaScript/CSS zamiast gating serwer-side. wysyłka the whole artykuł in the HTML i hiding it until login jest bypassable (disable JS i it’s readable), muddies Google’s ability to recognize the paywall, i makes screen czytelnicy przeczytaj “hidden” tekst aloud. Fix: confirm subscription/login status on the serwer i tylko wysyłać the pełny treść to authenticated użytkownicy — then dodaj dane strukturalne on top.
3. Assuming dowolny paywall counts as cloaking. Google’s spam polityka explicitly carves paywalls out of the cloaking definition, conditioned on letting Google see the pełny treść i following flexible-sampling guidance. Fix: don’t hide twój treść z Google out of cloaking fear — declare it z znaczniki, który jest the sanctioned mechanism.
4. Believing there’s a special “Googlebot Subscriber” crawler. Several niski-quality poradniki (prawdopodobny one propagating to others) twierdzenie you musi pozwalać a “Googlebot Subscriber” lub “Googlebot Registered User” crawler. No such użytkownik-agent exists — Google’s opublikowany crawler lista ma Googlebot, Googlebot-Image, Googlebot-Video, i Googlebot-News, i nothing subscriber-powiązany. Fix: ignore it; there’s no oddzielny crawler to pozwalać-lista.
5. Blocking private/login URLs z robots.txt.
A robots-blocked URL może nadal być zindeksowany as a bare, contentless URL — często worse
than a clean noindex, i Mueller says as much. Fix: używać noindex lub a
redirect dla private treść; reserve robots.txt dla crawl-budget control, nie
deindexing.
6. Citing an “80-word minimum lead-in” as Google polityka. ten figure circulates as if it’s official, ale it doesn’t trace to dowolny Google document. Google’s rzeczywisty quantified guidance jest o sampling frequency (6–10 artykuły/month), nie lead-in word count. Fix: treat dowolny word-count floor as an niezweryfikowany practitioner heuristic, nie polityka.
7. Forgetting że showing Google the pełny tekst needs a pamięć podręczna control.
The paywall znaczniki lets Googlebot see the pełny artykuł, ale a cached copy może leak
it to anyone who finds the pamięć podręczna. Fix: dodawać noarchive/nocache (lub the
X-Robots-Tag: noarchive header) if że’s a concern — it’s a oddzielny control z
the paywall znaczniki.
Snippets dla checking a paywall setup
Practical sprawdzenia dla whether twój gating i znaczniki actually działać. Swap
https://example.com/article i .paywall dla twój own.
1. robi the pełny artykuł ship in the HTML? (the JS-trap test)
If twój paid treść jest present in the raw serwer odpowiedź, it’s nie really gated — it’s just visually hidden. Fetch the HTML bez executing JavaScript i search dla a paid zdanie.
macOS / Linux (curl + grep)
# Fetch the raw HTML (no JS execution) and look for a line that should be gated.
curl -s "https://example.com/article" | grep -i "a sentence only subscribers should see"
# Empty result = the gated text isn't in the raw HTML (good, server-side gated).
# A match = the full content is shipping to everyone and merely hidden (the JS trap).porównywać co Googlebot vs. a logged-out użytkownik receives
# As a normal visitor:
curl -s "https://example.com/article" -o guest.html
# Emulating Googlebot's user-agent (only meaningful if you serve UA-based content):
curl -s -A "Mozilla/5.0 (compatible; Googlebot/2.1; +http://www.google.com/bot.html)" \
"https://example.com/article" -o googlebot.html
# Diff the visible article body — Googlebot should get the full text under flexible sampling.
diff <(grep -o '<p>.*</p>' guest.html) <(grep -o '<p>.*</p>' googlebot.html)2. Extract i sanity-sprawdź paywall dane strukturalne
Pull the JSON-LD bloki z a Chrome DevTools Console snippet. otwarty the artykuł, otwarty DevTools → Console, paste:
// Dump every JSON-LD block and flag paywall properties.
[...document.querySelectorAll('script[type="application/ld+json"]')]
.map(s => { try { return JSON.parse(s.textContent); } catch { return null; } })
.filter(Boolean)
.forEach(obj => {
const json = JSON.stringify(obj);
if (json.includes('isAccessibleForFree') || json.includes('cssSelector')) {
console.log('Paywall markup found:', obj);
} else {
console.log('JSON-LD (no paywall props):', obj['@type']);
}
});Confirm the cssSelector actually matches an element (class selectors tylko —
#id i complex selectors są unsupported):
// Paste your declared selector; it MUST match at least one element, and be a .class.
const sel = '.paywall';
console.log('Matches on page:', document.querySelectorAll(sel).length);
console.log('Is a class selector:', /^\.[\w-]+$/.test(sel)); // true = supported form3. Bookmarklet: jest ten strona marked as gated?
Drag-to-bookmark ten one-liner (lub paste in the address bar) to sprawdzenie dowolny artykuł
dla isAccessibleForFree: false bez opening DevTools:
javascript:(()=>{const b=[...document.querySelectorAll('script[type="application/ld+json"]')].map(s=>s.textContent).join('');alert(b.includes('"isAccessibleForFree":false')||b.includes('"isAccessibleForFree": false')?'Gated: isAccessibleForFree:false present':'No paywall markup found on this page');})();4. Confirm the cached copy isn’t leaking (noarchive sprawdzenie)
# Check for a noarchive directive in the meta robots tag or the X-Robots-Tag header.
curl -s "https://example.com/article" | grep -i 'name="robots"'
curl -sI "https://example.com/article" | grep -i 'x-robots-tag'
# You want "noarchive" (or nocache) present if you don't want a cached copy exposing gated text.po te pass, walidować the live URL in the
wyniki z elementami rozszerzonymi Test as Googlebot komputer stacjonarny
i smartphone — it’s the authoritative sprawdzenie że Google parses twój
isAccessibleForFree/cssSelector znaczniki.
Paywall symptoms, causes, i fixes
wyniki z elementami rozszerzonymi Test robi nie walidować the gated sekcja
Symptom: The live URL robi nie pokazywać usable paywalled-treść znaczniki, lub the reported
cssSelector robi nie zidentyfikuj gated treść.
prawdopodobny cause: isAccessibleForFree jest missing lub ustawić inconsistently; hasPart jest
malformed; the selector używa an ID lub complex selector zamiast a class; the HTML lacks
the declared class; lub Googlebot receives tylko the teaser i cannot inspect the pełny działać.
Fix i confirmation: używać isAccessibleForFree: false on the działać i każdy gated
WebPageElement, point każdy cssSelector at an rzeczywisty class such as .paywall, i
make the pełny subscriber-equivalent treść available to Googlebot poniżej flexible
sampling. Re-run the live URL in the wyniki z elementami rozszerzonymi Test as oba smartphone i komputer stacjonarny
until the znaczniki i gated sekcja są parsed as intended.
Logged-out źródło contains the pełny paid artykuł
Symptom: Disabling JavaScript, inspecting the HTML, lub używając a screen czytelnik exposes tekst że the visible paywall twierdzenia jest unavailable.
prawdopodobny cause: The serwer ships the pełny artykuł to everyone i JavaScript lub CSS merely hides it po the strona loads.
Fix i confirmation: Move entitlement checking to the serwer i wysyłać the pełny tekst tylko po login/subscription jest confirmed, podczas gdy continuing to serve Googlebot poniżej the declared flexible-sampling setup. Fetch the strona logged out z scripts disabled i confirm the gated body jest absent; then authenticate i confirm the pełny artykuł arrives.
wyniki wyszukiwania lead mainly to a bare login strona
Symptom: An incognito marka/service search surfaces a generic login strona, lub wiele private URLs collapse onto the same contentless login experience.
prawdopodobny cause: Private routes redirect to one generic strona z no service context, lub robots.txt bloki the private URLs podczas gdy nadal allowing bare URL indeksowanie.
Fix i confirmation: Give legitimate login destinations unique contextual copy. dla
genuinely private treść, używać authentication plus noindex lub a purposeful login
redirect zamiast robots.txt as an indeksowanie control. Repeat the incognito search i
inspect representative URLs to confirm the wynik jest informative i private URLs są
nie appearing as bare entries.
Google może rank tylko the teaser
Symptom: The strona jest zindeksowany ale appears relevant tylko to the lead-in, nie to the pełny artykuł’s subject.
prawdopodobny cause: Googlebot receives the same short teaser as an unauthenticated czytelnik, so the engine cannot zrozum gated body.
Fix i confirmation: Implement flexible sampling so verified Googlebot może crawl the same pełny treść a subscriber receives, declare the gated sekcja z dane strukturalne, i walidować the live strona. używać URL Inspection po recrawl to confirm Google może render the intended artykuł; ranking odzyskiwanie jest nie an immediate walidacja signal.
Paywall launch checklist
Sampling i access model
- Choose metering lub lead-in deliberately; robić nie inherit an arbitrary vendor domyślny.
- If używając a meter, test monthly sampling pierwszy i używać Google’s 6–10 free artykuły-per-month range as a starting point, nie a universal command.
- monitorować how często the paywall appears; Google says satisfaction degrades gdy it jest shown więcej niż 10% of the time.
- Googlebot może access the same pełny treść an entitled czytelnik receives poniżej the declared flexible-sampling setup.
znaczniki
- The top-level
Article,NewsArticle, lub otherCreativeWorkdeclaresisAccessibleForFree: falsegdy the działać jest gated. - każdy gated sekcja ma a
hasPartWebPageElementzisAccessibleForFree: false. - każdy
cssSelectorużywa an rzeczywisty class selector such as.paywall, nie an ID lub complex descendant selector. - Multiple gated sekcje są oddzielny, non-nested
hasPartentries. - Registration walls użyj same paywall znaczniki as paid access walls.
Delivery i prywatność
- Entitlement jest enforced serwer-side; logged-out HTML robi nie contain the hidden pełny artykuł dla JavaScript lub CSS to reveal.
- The mobilny odpowiedź follows the same poprawny gating i sampling behavior.
- Genuinely private URLs używać authentication i
noindexlub a login redirect, nie robots.txt as the prywatność mechanism. - Login strony obejmować użyteczny, service-specific context zamiast one generic strona duplicated w całym każdy route.
-
noarchive/nocachejest present gdzie cached copies musi nie expose gated tekst.
Pre-launch proof
- The live candidate validates in the wyniki z elementami rozszerzonymi Test as smartphone i komputer stacjonarny.
- A logged-out, JavaScript-disabled fetch robi nie reveal the pełny gated body.
- An authenticated session receives the complete artykuł.
- An incognito marka/service search robi nie reduce the witryna to a bare login wynik.
- Analytics records meter consumption i paywall exposure bez w tym private artykuł tekst in event payloads.
Prove the paywall jest declared i enforced
Paywalled structured-data test
- Test to run: Test the live URL in Google’s wyniki z elementami rozszerzonymi Test as smartphone i
komputer stacjonarny, inspecting
isAccessibleForFree,hasPart, i każdycssSelector. - Expected wynik: Google parses the gated
CreativeWorki każdy declared class maps to the intended gated sekcja podczas gdy Googlebot może access the pełny artykuł. - awaria interpretation: Missing właściwości, selector mismatches, nieprawidłowy nesting, lub a teaser-tylko Googlebot odpowiedź means the flexible-sampling declaration jest broken.
- monitorowanie window: Immediate po każdy template lub paywall-vendor deployment.
- Rollback trigger: The production template stops declaring lub exposing gated treść correctly w całym the artykuł ustawić i cannot być fixed przed broad rollout.
serwer-side entitlement test
- Test to run: Fetch the same artykuł logged out z JavaScript disabled, then fetch it in an authenticated entitled session; obejmować a screen-czytelnik sprawdzenie on the logged-out odpowiedź.
- Expected wynik: Logged-out użytkownicy receive tylko the intended sample i cannot find the gated body in HTML/DOM, podczas gdy entitled użytkownicy receive the complete artykuł.
- awaria interpretation: pełny tekst in the logged-out odpowiedź means the paywall tylko hides treść client-side; missing tekst po authentication means entitlement delivery jest failing.
- monitorowanie window: Immediate in staging i production po dowolny paywall JavaScript, template, pamięć podręczna, CDN, lub authentication change.
- Rollback trigger: Unauthenticated użytkownicy może pobierać the pełny paid artykuł, lub entitled czytelnicy broadly lose access po the change.
pamięć podręczna-control i private-URL test
- Test to run: Inspect the strona’s robots meta i X-Robots-znacznik dla
noarchivelubnocacheas required, then inspect representative genuinely private URLs dla auth inoindexbehavior. - Expected wynik: Cached-copy controls są present on gated artykuły gdzie intended; private URLs są protected i nie relying on robots.txt alone to zapobiegać indeksowanie.
- awaria interpretation: Missing pamięć podręczna directives create a copy-leak risk, podczas gdy a robots-tylko blok może leave a bare private URL eligible dla indeksowanie.
- monitorowanie window: Immediate po header, CDN, robots, lub authentication changes; recheck the affected templates po deployment.
- Rollback trigger: A deployment exposes private treść, removes access controls, lub broadly makes private URLs indeksowalny i cannot być corrected immediately.
Ongoing flexible-sampling metrics
Paywall display rate
- Metric: The percentage of eligible treść views in który the paywall jest shown.
- co it tells you: How restrictive the sampling model feels w całym visits; it jest the exposure mierzyć Google ties directly to użytkownik satisfaction.
- How to pull it: Divide serwer- lub paywall-platforma gate impressions by eligible artykuł views, segmented by użytkownik cohort, acquisition źródło, i urządzenie.
- Benchmark / realistic range: Google says general satisfaction degrades significantly gdy paywalls są shown więcej niż 10% of the time, generally exposing o 3% of the odbiorcy. Treat że as a caution ceiling i test wobec twój own subscribers, firma model, i artykuł mix.
- Cadence: Weekly dla abrupt konfiguracja changes i monthly dla the stable trend; ten jest a leading experience/monetization control.
Monthly free-artykuł allowance i consumption
- Metric: The configured monthly free-artykuł quota plus the distribution of how wiele free artykuły użytkownicy consume przed encountering the gate.
- co it tells you: Whether the meter gives czytelnicy enough sampling to zrozum produkt podczas gdy nadal reaching the subscription prompt.
- How to pull it: używać serwer-side meter lub paywall-platforma logs grouped by anonymous meter tożsamość i month; raport the configured quota alongside consumption percentiles.
- Benchmark / realistic range: Google recommends 10 artykuły per month as an exploration starting point i expects 6–10 per użytkownik per month dla najbardziej daily news publishers. ten jest a starting range, nie a mandate dla każdy publikacja.
- Cadence: Monthly, matching the recommended meter period; sprawdzenie po deliberate quota experiments zamiast reacting to daily noise.
Test yourself: Paywalls i SEO
Five quick questions on keeping gated treść indeksowalny bez cloaking. Pick an answer dla każdy, then sprawdzenie.
Dziennik zmian
Zaktualizowano 18 lip 2026.
Podsumowanie redakcyjne i zapisane szczegóły zmian.Szczegóły zmian
-
Szczegółowe uwagi dotyczące zmian są obecnie dostępne po angielsku.
-
Szczegółowe uwagi dotyczące zmian są obecnie dostępne po angielsku.
-
Szczegółowe uwagi dotyczące zmian są obecnie dostępne po angielsku.
Pełne porównanie jest niedostępne — dla tej wersji nie zarchiwizowano wcześniejszej migawki.