Hướng dẫn về Meta Robots Tag
Đó robots meta tag controls cách một trang là được lập chỉ mục và phân phối — mỗi directive, đó crawl-thì-obey rule, conflict resolution, và meta tag so với X-Robots-Tag.
Ngôn ngữ
1 tín hiệu bằng chứng trên trang này
- Công cụ trực tuyến liên quanHTTP Header Checker
Đó robots meta tag — <meta name="robots" content="noindex"> trong đó <head> — tells các công cụ tìm kiếm cách chỉ mục và serve một single trang. Đó rule đó breaks mọi thứ: đây là crawl-thì-obey, so một trang blocked trong robots.txt là không bao giờ fetched và của nó noindex là không bao giờ seen. Với không tag, đó default là chỉ mục, follow. Conflicting rules resolve để đó hầu hết restrictive; cho một googlebot-named tag so với đó generic robots tag, Googlebot takes đó sum of đó negative rules. Đó tag là HTML-chỉ — dùng đó X-Robots-Tag header cho PDFs, images, và other non-HTML.
TL;DR — Đó robots meta tag là một line of HTML bạn put trong một trang
<head>để tell các công cụ tìm kiếm cách xử lý đó một trang — hầu hết thường<meta name="robots" content="noindex">để giữ điều này out of tìm kiếm. Đó catch đó trips mọi người lên: Google có để là able để crawl đó trang để đọc đó tag. Nếu bạn cũng block đó trang trongrobots.txt, Google không bao giờ sees đó tag, và đó trang có thể stay trong tìm kiếm. Với không tag tại all, đó default là “index it and follow the links.” (bản dịch) «chỉ mục điều này và follow đó links.»
Điều gì robots meta tag là
robots meta tag là nhỏ HTML element đó sits trong <head> của trang:
<meta name="robots" content="noindex">Điều này tells các công cụ tìm kiếm cách treat này cụ thể trang — liệu để cho thấy điều này
trong kết quả, liệu để follow đó links on điều này, liệu để cho thấy một snippet, và so
on. Đó name="robots" part có nghĩa là “all search engines that read this tag.” (bản dịch) «all các công cụ tìm kiếm đó đọc này tag.» Bạn
có thể đổi trong một crawler name, như name="googlebot", để talk để chỉ một engine.
nếu có không robots meta tag trên một trang, default là index, follow — hiển thị
nó trong tìm kiếm và follow của nó links. So bạn chỉ cần tag Khi bạn muốn để
thay đổi đó default. Evidence for this claim For Google, the default robots meta behavior is index, follow when no restrictive rule is present. Scope: Google-supported robots meta rules; other crawlers publish their own support and defaults. Confidence: high · Verified: Google Search Central: Robots meta tag specifications
một điều để nhận right: không block trang bạn’re trying để noindex
Đây là mistake I see phần lớn thường. Mọi người muốn trang out của Google, so họ làm cả hai điều tại sau khi:
- Thêm
noindexđể trang, và - Block trang trong
robots.txt.
Đó second step defeats đó đầu tiên. Blocking một URL trong robots.txt tells Google
“don’t even fetch this page.” (bản dịch) «không ngay cả fetch này trang.» So Google không bao giờ downloads điều này, không bao giờ đọc đó
noindex, và đó trang có thể stay trong đó chỉ mục. Đó cách sửa là để leave đó trang
crawlable và let Google đọc đó noindex. Evidence for this claim Google can read and follow page-level robots rules only when it is allowed to access the page. Scope: Google-supported robots meta and X-Robots-Tag rules; robots.txt blocking can prevent rule discovery. Confidence: high · Verified: Google Search Central: Robots meta tag specifications
Think của nó as hai khác jobs:
robots.txtcontrols crawling — liệu bots fetch trang tại all.- ** robots meta tag** controls lập chỉ mục và serving — Điều gì happens sau khi trang là fetched.
họ’re không interchangeable, và bạn không muốn cả hai on giống nhau URL Khi của bạn goal là để xóa nó từ tìm kiếm.
The same page contains a meta robots noindex directive. In the first path, crawling is allowed, so the crawler fetches the page, reads noindex, and can remove the URL from results after processing. In the second path, robots.txt blocks crawling, so the crawler cannot fetch the page or see noindex, and the URL may remain in results. Crawl control and index control are separate jobs.
© Patrick Stox LLC · CC BY 4.0 ·
phổ biến directives
một vài bạn’ll thực ra sử dụng:
noindex— giữ điều này trang out của kết quả tìm kiếm.nofollow— không follow links on điều này trang. (khác từ puttingrel="nofollow"on single link — điều này một áp dụng để mỗi link on trang.)none— shorthand chonoindex, nofollow.nosnippet— không hiển thị text snippet cho điều này trang trong kết quả.
Bạn có thể combine them với comma: <meta name="robots" content="noindex, nofollow">.
một vài honest gotchas
noindexkhông làm trang riêng tư. trang là vẫn công khai và crawlable — nó chỉ sẽ không hiển thị trong tìm kiếm. cho thực quyền riêng tư, sử dụng login.noindexkhông save ngân sách crawl. Google vẫn có để fetch trang để see tag.- cho PDF hoặc image, Bạn có thể’t thêm
<meta>tag — có không HTML<head>. đó Điều gì X-Robots-Tag HTTP header là cho (nhiều hơn on đó trong Nâng cao version). - có không có mốc thời gian cố định cho
noindexed trang để drop out của Google. Google nói nó có thể take months cho thấp hơn-priority trang để nhận recrawled và processed — không promise client “2-4 weeks.”
Muốn đầy đủ directive list, Cách Google resolves conflicting rules,
googlebot-so với-robots trường hợp biên, và X-Robots-Tag header trong detail? Chuyển
để Nâng cao tab.
Tự kiểm tra: robots meta directives
chọn control by outcome
Which robots control should I use?
Robots controls đó fail của họ dự kiến job
- Combining
noindexvới robots.txt block. block ngăn crawler từ seeing removal directive. - sử dụng
disallowas deindexing bảo đảm. nó dừng fetching, nhưng linked URL có thể vẫn là được lập chỉ mục không có nội dung trang. - Assuming thân phản hồi-placed robots meta tag là silently đã bỏ qua. Google nói nó
sẽ respect robots meta tag trong thân phản hồi, so template đó injects nó ở đó
vẫn hoạt động cho Google — nhưng stick để
<head>placement anyway, since nó chỉ placement khác các crawler và các validator là guaranteed để expect. - sử dụng crawler-name token Google không đọc (e.g.
name="bingbot") và assuming nó cũng scopes rule cho Google. Google recognizes chỉgooglebotvàgooglebot-news; bất kỳ khác name giá trị là đã bỏ qua by Google hoàn toàn. - sử dụng HTML meta tag cho PDF. phục vụ
X-Robots-Tagtrong HTTP phản hồi cho non-HTML files. - Treating
nofollowasnoindex. Link xử lý không xóa hiện tại trang từ kết quả. - Xuất bản generic và crawler-cụ thể tags không có resolving sum. Audit
mỗi
robots,googlebot, và header directive together; restrictive rule có thể survive trong thứ hai location. - sử dụng robots rules cho secrecy. Anyone có thể yêu cầu công khai URL hoặc đọc
robots.txt; confidential nội dung cần authentication.
Deploy noindex safely
- Confirmed removal từ tìm kiếm—không crawl reduction, canonical consolidation, hoặc access control—là dự kiến outcome.
- được sử dụng
<meta name="robots" content="noindex">cho HTML hoặcX-Robots-Tag: noindexphản hồi header cho non-HTML tài nguyên. - Kept URL crawlable và accessible để đích crawler.
- Đã xóa contradictory template, CMS, CDN, và crawler-cụ thể directives.
- Checked được kết xuất head và trực tiếp phản hồi, không chỉ nguồn template.
- Tested cả hai canonical URL và có ý nghĩa variants hoặc các chuyển hướng.
- Recorded affected URL đặt và deployment time cho sau đó so sánh.
- Verified rule trong URL Inspection và monitored trang lập chỉ mục báo cáo sau khi recrawl.
- được sử dụng authentication thay vì nếu nội dung phải là riêng tư.
Inspect meta, các header, và crawl access
Replace URL trong điều này shell audit:
url="https://example.com/page/"
curl -sSI "$url" | grep -iE '^(HTTP/|x-robots-tag:|location:)'
curl -sS "$url" | grep -oiE '<meta[^>]+name=["'"'](robots|googlebot)["'"'][^>]*>'
curl -sS "https://example.com/robots.txt"Trong DevTools Console, inventory mỗi parsed robots tag thay vì stopping tại đầu tiên match:
console.table([...document.querySelectorAll('meta[name]')]
.filter((meta) => /^(robots|googlebot|bingbot)$/i.test(meta.name))
.map((meta) => ({ crawler: meta.name, content: meta.content })));Các header không phải visible trong DOM, so sánh console kết quả với thực tế phản hồi. cũng follow các chuyển hướng: directive on intermediate phản hồi không prove đích phục vụ giống nhau rule.
Kiểm thử mỗi layer riêng
- HTTP Header Checker — inspect status, các chuyển hướng,
và mỗi trực tiếp
X-Robots-Tagheader, including cho PDFs. - Robots.txt Tester — kiểm tra liệu đích crawler là được phép để fetch URL và do đó able để discover robots directive.
sử dụng cả hai Khi diagnosing stubborn được lập chỉ mục URL. đúng noindex phản hồi không phải actionable nếu robots.txt ngăn fetch, và được phép crawl không prove trang thực ra phục vụ noindex.
Prove directive hoạt động sau khi deployment
Kiểm thử 1 — Trực tiếp phân phối
- Hypothesis: mỗi dự kiến URL phục vụ một effective
noindexrule. - Phương thức: Sample URL đặt, follow các chuyển hướng, và inspect được kết xuất meta plus phản hồi các header.
- Truyền condition: cuối phản hồi là crawlable và exposes
noindexđể dự kiến crawler với không conflicting phân phối path. - Fail condition: chuyển hướng, CDN header, hoặc crawler-cụ thể tag thay đổi nó.
- tiếp theo hành động: khắc phục responsible layer và repeat giống nhau sample.
Kiểm thử 2 — Tìm kiếm-engine processing
- Hypothesis: Google có thể crawl trang và có processed removal rule.
- Phương thức: Chạy URL Inspection trực tiếp kiểm thử, sau đó monitor trang lập chỉ mục báo cáo sau khi recrawl.
- Truyền condition: trực tiếp kiểm thử detects noindex và URL becomes không được lập chỉ mục cho đó reason.
- Fail condition: URL là blocked, directive là absent, hoặc old chỉ mục state persists không có new crawl.
- tiếp theo hành động: Restore crawl access hoặc yêu cầu recrawl; không thêm disallow.
Đo lường intent, không universal benchmark
Dự kiến-noindex coverage
URLs serving the intended noindex ÷ URLs in the approved noindex set
Xây dựng denominator từ của bạn own removal inventory, sau đó so sánh crawl trước khi và sau khi deployment. đích là hoàn tất coverage của đó approved đặt—không ngành percentage.
Unintended chỉ mục-control conflicts
Track được tính của dự kiến-noindex các URL đó là vẫn được lập chỉ mục, robots.txt-blocked, hoặc serving contradictory meta/header rules. Segment by template, CDN rule, và file loại so một implementation fault không hide bên trong sitewide total.
sử dụng Google Search Console’s trang lập chỉ mục báo cáo và URL Inspection as processing evidence, trong khi recognizing đó deindexing phụ thuộc vào recrawl. Bảo toàn pre-thay đổi count và deployment date; nếu không falling total có không reliable baseline.
Evidence for this claim Google can read and follow page-level robots rules only when it is allowed to access the page. Scope: Google-supported robots meta and X-Robots-Tag rules; robots.txt blocking can prevent rule discovery. Confidence: high · Verified: Google Search Central: Robots meta tag specificationsTL;DR —
<meta name="robots" content="…">trong đó<head>controls cách một single trang là được lập chỉ mục và phân phối; với không tag đó default làindex, follow. đây là crawl-thì-obey: “these settings can be read and followed only if crawlers are allowed to access the pages” (bản dịch) «những settings có thể là đọc và followed chỉ nếu các crawler là được phép để access đó các trang» — so mộtrobots.txt-blocked URL là không bao giờ fetched và của nónoindexlà không bao giờ seen (I có đầu tiên-party dữ liệu on đó flip side of này). Conflicting rules resolve để đó hầu hết restrictive; trên mộtgooglebottag và đó genericrobotstag, Googlebot takes đó sum of đó negative rules. Đó tag là HTML-chỉ — dùng đó X-Robots-Tag header cho non-HTML và tại quy mô. Google đọc chỉ hai crawler-named tokens —googlebotvàgooglebot-news— và bỏ qua mỗi other giá trị, including other engines’ tokens nhưbingbot.
Điều gì nó là, nơi nó goes, và default
Đó robots meta tag cho phép bạn, trong Google words, “use a granular, page-specific
approach to controlling how an individual HTML page should be indexed and served
to users in Google Search results.” (bản dịch) «dùng một granular, trang-cụ thể approach để controlling cách an riêng lẻ HTML trang nên là được lập chỉ mục và phân phối để người dùng trong Google Search kết quả.» Đó conventional, portable place cho điều này là
đó <head>:
<meta name="robots" content="noindex, nofollow">Head placement là đó authoring convention mỗi engine expects, nhưng Google là
rõ ràng đó điều này không một hard requirement cho Google Search cụ thể: “Google
Search doesn’t enforce placement of meta robots in the HTML head and will respect
robots meta tags in the body section of an HTML document as well.” (bản dịch) «Google Search không enforce placement of meta robots trong đó HTML head và sẽ respect robots meta tags trong đó thân phản hồi section of an HTML document as well.» Treat đó as
tolerance cho Google, không portable advice — vẫn tác giả điều này trong đó <head> so
mỗi crawler và validator đó expects tiêu chuẩn placement đọc điều này correctly.
Đó name thuộc tính là đó audience, và này là nơi hầu hết các hướng dẫn overstate
Google hỗ trợ. name="robots" addresses mỗi crawler đó đọc đó tag.
Beyond đó, Google hỗ trợ chính xác hai crawler-named tokens, và bỏ qua mỗi
other giá trị: “Google supports two user agent tokens in the robots meta tag;
other values are ignored: googlebot for all text results, and googlebot-news
for news results.” (bản dịch) «Google hỗ trợ hai người dùng agent tokens trong đó robots meta tag; other các giá trị là đã bỏ qua: googlebot cho all text kết quả, và googlebot-news cho news kết quả.» MỘT tag named name="bingbot" không một được ghi lại Google
control — Bing đọc của nó own token on của nó own terms, nhưng Google skips bất kỳ name
giá trị điều này không recognize. Cả hai đó name và content các thuộc tính là
case-insensitive để Google, và so là X-Robots-Tag header names và các giá trị.
Khi không robots meta tag là present, đó default là
index, follow (đó all rule, mà Google notes “has no effect if explicitly
listed” (bản dịch) «có không effect nếu explicitly listed»). Evidence for this claim For Google, the default robots meta behavior is index, follow when no restrictive rule is present. Scope: Google-supported robots meta rules; other crawlers publish their own support and defaults. Confidence: high · Verified: Google Search Central: Robots meta tag specifications Bạn chỉ cần đó tag để thay đổi đó default.
Bạn combine rules hai ways: comma-separated trong một tag (noindex, nofollow) hoặc as
multiple <meta> tags. Google: bạn có thể “create a multi-rule instruction by
combining robots meta tag rules with commas or by using multiple meta tags.” (bản dịch) «tạo một multi-rule instruction by combining robots meta tag rules với commas hoặc by dùng multiple meta tags.»
rule đó breaks mọi thứ: Google phải crawl trang để see tag
Đây là toàn bộ bài viết. robots meta tag là crawl-sau đó-obey. Google có để fetch trang để đọc tag — so bất cứ điều gì đó dừng fetch dừng tag từ bao giờ là applied. Straight từ spec:
Evidence for this claim Google can read and follow page-level robots rules only when it is allowed to access the page. Scope: Google-supported robots meta and X-Robots-Tag rules; robots.txt blocking can prevent rule discovery. Confidence: high · Verified: Google Search Central: Robots meta tag specifications“Keep in mind that these settings can be read and followed only if crawlers are allowed to access the pages that include these settings.” (bản dịch) «Hãy nhớ rằng đó những settings có thể là đọc và followed chỉ nếu các crawler là được phép để access đó các trang đó bao gồm những settings.»
và consequence, spelled out:
“If a page is disallowed from crawling through the robots.txt file, then any information about indexing or serving rules will not be found and will therefore be ignored.” (bản dịch) «Nếu một trang là disallowed từ crawling qua đó robots.txt file, thì bất kỳ information về lập chỉ mục hoặc serving rules sẽ không là được tìm thấy và sẽ do đó là đã bỏ qua.»
So đó classic mistake — Disallow trong robots.txt plus noindex on đó
giống nhau URL — silently defeats đó noindex. Google không bao giờ crawl đó trang, không bao giờ
sees đó tag, và đó URL có thể linger trong đó chỉ mục (thường as một bare, snippet-ít hơn
kết quả nếu điều gì đó links để điều này). Google companion “Block Search Indexing” (bản dịch) «Block Tìm kiếm Lập chỉ mục»
doc says cùng một điều trong plainer language: cho đó noindex rule để hoạt động, đó
trang “must not be blocked by a robots.txt file… If the page is blocked by a
robots.txt file or the crawler can’t access the page, the crawler will never see
the noindex rule, and the page can still appear in search results.” (bản dịch) «không được là blocked by một robots.txt file… Nếu đó trang là blocked by một robots.txt file hoặc đó crawler không thể access đó trang, đó crawler sẽ không bao giờ see đó noindex rule, và đó trang có thể vẫn xuất hiện trong kết quả tìm kiếm.»
I’ve watched đó flip side of này mechanism happen với real dữ liệu. Trong my
thử nghiệm Đó Story of Blocking 2 Cao-Xếp hạng Các trang Với Robots.txt,
I có chủ ý blocked hai of của chúng ta xếp hạng các trang trong robots.txt. Vì Google
có thể không lâu hơn crawl them, điều này không thể refresh bất cứ điều gì về them — và đó
các trang mostly kept xếp hạng: “We lost a position here or there and all of the
featured snippets for the pages.” (bản dịch) «We lost một position ở đây hoặc ở đó và all of đó featured snippets cho đó các trang.» My takeaway: “Accidentally blocking pages
(that Google already ranks) from being crawled using robots.txt probably isn’t
going to have much impact on your rankings, and they will likely still show in the
search results.” (bản dịch) «Accidentally blocking các trang (đó Google đã ranks) từ đang được crawl dùng robots.txt probably không going để có nhiều impact on của bạn thứ hạng, và they sẽ có khả năng vẫn cho thấy trong đó kết quả tìm kiếm.» đó là đó giống nhau coin as đó noindex vấn đề — một blocked URL là
frozen. Block ≠ xóa. Nếu bạn thực ra muốn một trang đã biến mất, bạn cần một
crawlable noindex, mà là đó entire point of này tag.
Này là đó line I’ve drawn publicly on nơi mỗi tool belongs. Asked liệu
Google nên thêm noindex hỗ trợ để robots.txt, I đã nói: “Google was clear
they want robots.txt for crawl control only.” (bản dịch) «Google đã là clear they muốn robots.txt cho crawl control chỉ.» Crawling là robots.txt job;
lập chỉ mục là đó meta tag (hoặc đó header). They không overlap, và đó
noindex directive trong robots.txt đã là không bao giờ officially supported — Google dropped
phân tích cú pháp of điều này on September 1, 2019.
mỗi robots meta directive ( reference)
Google supported các giá trị, với verbatim các mô tả từ spec:
lập chỉ mục
all— “There are no restrictions for indexing or serving. This rule is the default value and has no effect if explicitly listed.” (bản dịch) «Có không restrictions cho lập chỉ mục hoặc serving. Này rule là đó default giá trị và có không effect nếu explicitly listed.»noindex— “Do not show this page, media, or resource in search results.” (bản dịch) «Không cho thấy này trang, media, hoặc tài nguyên trong kết quả tìm kiếm.»none— “Equivalent tonoindex, nofollow.” (bản dịch) «Tương đương đểnoindex, nofollow.»indexifembedded— “Google is allowed to index the content of a page if it’s embedded in another page through iframes or similar HTML tags, in spite of anoindexrule.” (bản dịch) «Google là được phép để chỉ mục đó nội dung of một trang nếu đây là embedded trong một sản phẩm khác trang qua iframes hoặc similar HTML tags, trong spite of mộtnoindexrule.» (Đó một directive đó overrides mộtnoindex, cho embedded nội dung.)
Links
nofollow— “Do not follow the links on this page.” (bản dịch) «Không follow đó links on này trang.» Này là trang-cấp độ — khác nhau phạm vi từ một theo-linkrel="nofollow", mà áp dụng để một link.
Serving và snippets
nosnippet— “Do not show a text snippet or video preview in the search results for this page.” (bản dịch) «Không cho thấy một text snippet hoặc video preview trong đó kết quả tìm kiếm cho này trang.» Này phạm vi là rộng hơn đó classic text snippet: Google says điều này “applies to all forms of search results (at Google: web search, Google Images, Discover, AI Overviews, AI Mode) and will also prevent the content from being used as a direct input for AI Overviews and AI Mode.” (bản dịch) «áp dụng để all forms of kết quả tìm kiếm (tại Google: web tìm kiếm, Google Images, Discover, AI Overviews, AI Chế độ) và sẽ cũng ngăn đó nội dung từ đang dùng as một trực tiếp input cho AI Overviews và AI Chế độ.»max-snippet:[number]— “Use a maximum of [number] characters as a textual snippet for this search result.” (bản dịch) «Dùng một maximum of [number] characters as một textual snippet cho này tìm kiếm kết quả.» Giống nhau broadened phạm vi asnosnippet: điều này “applies to all forms of search results (such as Google web search, Google Images, Discover, Assistant, AI Overviews, AI Mode) and will also limit how much of the content may be used as a direct input for AI Overviews and AI Mode.” (bản dịch) «áp dụng để all forms of kết quả tìm kiếm (such as Google web tìm kiếm, Google Images, Discover, Assistant, AI Overviews, AI Chế độ) và sẽ cũng limit cách nhiều of đó nội dung có thể là dùng as một trực tiếp input cho AI Overviews và AI Chế độ.» đó là một trực tiếp-input eligibility control cho Google own AI tìm kiếm features — điều này không phải một chung AI-training opt-out. Giữ nội dung của bạn out of model training (e.g. Google-Extended) hoặc out of Tìm kiếm tách biệt generative-AI thuộc tính-cấp độ control trong Search Console là khác nhau các hệ thống với khác nhau scopes; không treatnosnippet/max-snippetas covering either.max-image-preview:[setting]— “Set the maximum size of an image preview for this page in search results.” (bản dịch) «Set đó maximum size of an image preview cho này trang trong kết quả tìm kiếm.» Settings:none,standard, hoặclarge(“A larger image preview, up to the width of the viewport, may be shown.” (bản dịch) «MỘT lớn hơn image preview, lên để đó width of đó viewport, có thể là shown.»).max-video-preview:[number]— “Use a maximum of [number] seconds as a video snippet for videos on this page in search results.” (bản dịch) «Dùng một maximum of [number] seconds as một video snippet cho videos on này trang trong kết quả tìm kiếm.»notranslate— “Don’t offer translation of this page in search results.” (bản dịch) «không offer translation of này trang trong kết quả tìm kiếm.»noimageindex— “Do not index images on this page.” (bản dịch) «Không chỉ mục images on này trang.»unavailable_after:[date/time]— “Do not show this page in search results after the specified date/time.” (bản dịch) «Không cho thấy này trang trong kết quả tìm kiếm sau đó specified date/time.»
Lịch sử — không lâu hơn active Google controls
một vài directives đó vẫn circulate trong older các hướng dẫn là ones Google nói nó không lâu hơn dùng. không thêm những điều này expecting them để làm bất cứ điều gì:
noarchive— “Thenoarchiverule is no longer used by Google Search to control whether a cached link is shown in search results, as the cached link feature no longer exists.” (bản dịch) «Đónoarchiverule là không lâu hơn dùng by Google Search để control liệu một được lưu đệm link là shown trong kết quả tìm kiếm, as đó được lưu đệm link feature không lâu hơn tồn tại.»nocache(một synonym some engines dùng chonoarchive) — “Thenocacherule isn’t used by Google Search.” (bản dịch) «Đónocacherule không dùng by Google Search.»nositelinkssearchbox— “Thenositelinkssearchboxrule is no longer used by Google Search to control whether the sitelink search box is shown for a given page, as the feature no longer exists.” (bản dịch) «Đónositelinkssearchboxrule là không lâu hơn dùng by Google Search để control liệu đó sitelink tìm kiếm box là shown cho một được cho trang, as đó feature không lâu hơn tồn tại.»
Paragraph-cấp độ (không trong meta tag)
có một sub-trang control: đó data-nosnippet thuộc tính. Google: bạn có thể
“designate textual parts of an HTML page not to be used as a snippet… on
span, div, and section elements.” (bản dịch) «designate textual parts of an HTML trang không để là dùng as một snippet… on span, div, và section elements.» Mọi thứ trong đó meta tag là trang-wide;
nosnippet / dữ liệu-nosnippet là cách bạn giữ một passage out of đó snippet
không có touching đó rest.
Combining directives và resolving conflicts
Hai rules govern Điều gì happens Khi directives collide.
1. Đó hơn restrictive rule wins. “In the case of conflicting robots rules,
the more restrictive rule applies. For example, if a page has both max-snippet:50
and nosnippet rules, the nosnippet rule will apply.” (bản dịch) «Trong đó case of conflicting robots rules, đó hơn restrictive rule áp dụng. Ví dụ, nếu một trang có cả hai max-snippet:50 và nosnippet rules, đó nosnippet rule sẽ apply.» nosnippet là stricter
hơn một 50-character cap, so nosnippet là điều gì bạn nhận.
2. googlebot so với robots — đó sum of đó negative rules. Này là đó một
hầu hết các hướng dẫn nhận sai. MỘT googlebot-named tag làm không đơn giản replace đó
generic robots tag — cho đó overlap, Googlebot takes đó union of đó
restrictions. Google: “For situations where multiple crawlers are specified
along with different rules, the search engine will use the sum of the negative
rules.” (bản dịch) «Cho situations nơi multiple các crawler là specified along với khác nhau rules, đó công cụ tìm kiếm sẽ dùng đó sum of đó negative rules.» Của họ worked ví dụ:
<meta name="robots" content="nofollow">
<meta name="googlebot" content="noindex">“The page containing these meta tags will be interpreted as having a noindex, nofollow rule when crawled by Googlebot.” (bản dịch) «Đó trang containing những meta tags sẽ là interpreted as có một noindex, nofollow rule khi được crawl by Googlebot.» Đó nofollow từ robots plus
đó noindex từ googlebot thêm lên để noindex, nofollow cho Googlebot. (Nơi
một googlebot tag và một robots tag set đó giống nhau directive differently, đó
crawler-named một là đó một đó áp dụng để đó crawler.)
Meta robots tag so với X-Robots-Tag ( HTTP header)
Đó robots meta tag là HTML-chỉ — điều này cần một <head>. Cho bất cứ điều gì đó
không HTML, bạn dùng đó X-Robots-Tag, mà delivers đó chính xác giống nhau rule
vocabulary trong đó HTTP header phản hồi. Google: “The X-Robots-Tag can be used
as an element of the HTTP header response for a given URL. Any rule that can be
used in a robots meta tag can also be specified as an X-Robots-Tag.” (bản dịch) «Đó X-Robots-Tag có thể là dùng as an element of đó HTTP header phản hồi cho một được cho URL. Bất kỳ rule đó có thể là dùng trong một robots meta tag có thể cũng là specified as an X-Robots-Tag.» Và đó
reason điều này tồn tại: “You can use the X-Robots-Tag for non-HTML files like image
files where the usage of robots meta tags in HTML is not possible.” (bản dịch) «Bạn có thể dùng đó X-Robots-Tag cho non-HTML files như image files nơi đó usage of robots meta tags trong HTML không phải có thể.»
So:
- PDF, image, hoặc khác non-HTML? Bạn có thể’t thêm
<meta>tag — sử dụng header, e.g.X-Robots-Tag: noindex. - Toàn bộ directories hoặc patterns? header là đặt tại máy chủ/CDN cấp độ, so nó scales để đểàn bộ paths trong một config rule.
- Một engine? header có thể đích crawler cũng:
X-Robots-Tag: googlebot: noindex, nofollow, và multiple X-Robots-Tag các header có thể là combined trong một phản hồi.
giống nhau rules, hai phân phối mechanisms: meta tag cho HTML các trang, header cho mọi thứ khác và cho quy mô.
Mà directives Bing và khác engines hỗ trợ
không assume directive đặt là universal — nó không phải. Bing hỗ trợ cốt lõi
lập chỉ mục và serving rules — noindex, nofollow, noarchive (với nocache as
của nó synonym), và nosnippet — và nó honors X-Robots-Tag cho non-HTML
các tài nguyên. nhưng Bing làm không hỗ trợ none shorthand, so cho cross-engine
safety, ghi noindex, nofollow out explicitly thay vì relying on none.
snippet- và preview-control family — max-snippet, max-image-preview,
max-video-preview — along với noimageindex, notranslate,
indexifembedded, và unavailable_after, là effectively Google-chỉ. Khi trong
doubt, spell directives out và treat max-* controls as Google features.
phổ biến mistakes (và các cách sửa)
Disallow+noindexon đó giống nhau URL. Đónoindexlà không bao giờ seen. Cách sửa: leave đó trang crawlable; giữ chỉ đónoindex.noindexplus mộtrel=canonicalpointing elsewhere. Conflicting các tín hiệu — bạn là telling Google cả hai “drop this page” (bản dịch) «drop này trang» và “consolidate it into another one.” (bản dịch) «consolidate điều này vào một sản phẩm khác một.» Pick một. (Hơn trong canonicalization.)- MỘT staging-wide
noindexshipped để production. Catastrophic, sitewide deindex. Kiểm tra trước launch. - MỘT
noindexinjected chỉ by client-side JavaScript. Google có để render đó trang để see điều này, và nếu đó được kết xuất HTML differs từ điều gì bạn expect, behavior differs cũng. Ưu tiên đó tag trong đó thô HTML hoặc đó header. (See kết xuất.) - Expecting Bing để honor Google-chỉ directives (
none, đómax-*family). - Expecting một robots directive để làm một job điều này không own. MỘT
noindexhoặcnosnippetrule không by itself bảo đảm crawl-budget savings, secrecy, một xếp hạng thay đổi, giống hệt behavior trên các công cụ tìm kiếm, một cụ thể removal timeline, hoặc exclusion từ mỗi AI/tìm kiếm surface — mỗi of những outcomes belongs để một khác nhau control (authentication cho secrecy,robots.txtcho crawl, mỗi engine own tài liệu cho parity, Search Console hoặc Google-Extended cho AI-cụ thể scopes). Google cho không fixed timeframe cho khi mộtnoindexed trang thực ra drops out — điều này phụ thuộc vào recrawl priority và “may take months” (bản dịch) «có thể take months» cho một thấp hơn-importance trang.
cho nơi điều này sits trong bigger picture: robots.txt và crawling là crawl-control side; noindex và lập chỉ mục là chỉ mục-control side; và nosnippet / dữ liệu-nosnippet, max-snippet, và max-image-preview là serving controls bạn reach cho Khi bạn muốn trang được lập chỉ mục nhưng muốn để shape Cách nó xuất hiện. X-Robots-Tag là điều này giống nhau tag HTTP-header tương đương cho non-HTML.
AI summary
condensed take on Nâng cao version:
- Đó robots meta tag —
<meta name="robots" content="…">trong đó<head>— controls cách một single HTML trang là được lập chỉ mục và phân phối. Với không tag, đó default làindex, follow. - đây là crawl-thì-obey. Google phải crawl trang để đọc đó tag: “these
settings can be read and followed only if crawlers are allowed to access the
pages.” (bản dịch) «những settings có thể là đọc và followed chỉ nếu các crawler là được phép để access đó các trang.» So một robots.txt-blocked URL không bao giờ nhận của nó
noindexseen — đó classic mistake. Patrick blocked-các trang thử nghiệm proves đó mechanism: blocked các trang stayed được lập chỉ mục và mostly kept xếp hạng. Block ≠ xóa. - robots.txt = crawl control; đó meta tag (hoặc X-Robots-Tag) = chỉ mục/serve
control. Không interchangeable.
noindextrong robots.txt đã là dropped Sept 1, 2019. - Directives: lập chỉ mục (
noindex,none,indexifembedded), links (nofollow, trang-wide), serving (nosnippet,max-snippet,max-image-preview,max-video-preview,notranslate,noimageindex,unavailable_after), plus đó paragraph-cấp độdata-nosnippetthuộc tính.noarchive,nocache, vànositelinkssearchboxlà lịch sử — Google says điều này không lâu hơn dùng them. nosnippet/max-snippetcũng gate AI Overviews và AI Chế độ — Google says they control liệu đó trang nội dung có thể là dùng as một trực tiếp input cho những features, không chỉ đó classic text snippet. đó là không một chung AI-training opt-out; Google-Extended và Tìm kiếm tách biệt generative-AI thuộc tính control là khác nhau các hệ thống.- Conflicts resolve để đó hầu hết restrictive (
nosnippetbeatsmax-snippet:50). Trên mộtgooglebottag và đó genericrobotstag, Googlebot takes đó sum of đó negative rules —robots: nofollow+googlebot: noindex⇒noindex, nofollow. Google recognizes chỉ đógooglebot/googlebot-newsname tokens; other các giá trị (nhưbingbot) là đã bỏ qua by Google. - Meta tag là conventionally
<head>-chỉ (Google cũng tolerates thân phản hồi placement, nhưng đó là Google-cụ thể, không portable); đó X-Robots-Tag carries đó giống nhau rules trong đó HTTP header cho PDFs, images, non-HTML, và toàn bộ directories. - Cross-engine: Bing hỗ trợ
noindex/nofollow/noarchive(nocache)/nosnippetnhưng khôngnone; đómax-*family là effectively Google-chỉ — ghi directives out explicitly. - Không bảo đảm: một robots directive alone không promise crawl-budget
savings, secrecy, một xếp hạng thay đổi, cross-engine parity, hoặc một fixed removal
timeline — Google cho không set timeframe cho
noindexđể take effect.
Tài liệu chính thức
Chính-nguồn tài liệu từ các công cụ tìm kiếm.
- Robots meta tag, dữ liệu-nosnippet, và X-Robots-Tag các đặc tả — có thẩm quyền spec: mỗi directive, combining rules, conflict resolution,
googlebot-so với-robotsunion, và X-Robots-Tag header. Bắt đầu ở đây. - Block Tìm kiếm lập chỉ mục với noindex — Cách
noindexhoạt động, hai phân phối mechanisms (meta tag và header), và crawl dependency trong đơn giản language. - Introduction để robots.txt — crawl-control đối tác tương ứng, so bạn không confuse hai jobs.
- Crawling và lập chỉ mục — hub cho robots, sitemaps, canonicalization, và crawl controls.
Bing / Microsoft
- Robots meta tags và các thuộc tính đó Bing hỗ trợ — Bing supported directives (xác nhận chính xác list on trực tiếp trang; nó JavaScript-được kết xuất).
Reference
- MDN —
<meta name="robots">— cross-engine directive reference, including mà engines sử dụngnoarchive/nocache.
Quotes từ nguồn
On—record statements từ Google. mỗi link là deep link đó jumps để quoted passage on nguồn trang.
Google — Điều gì tag là và default
- “The robots
metatag lets you use a granular, page-specific approach to controlling how an individual HTML page should be indexed and served to users in Google Search results.” (bản dịch) «Đó robotsmetatag cho phép bạn dùng một granular, trang-cụ thể approach để controlling cách an riêng lẻ HTML trang nên là được lập chỉ mục và phân phối để người dùng trong Google Search kết quả.» — Google Search Central tài liệu. Nhảy đến trích dẫn
Google — crawl-sau đó-obey dependency
- “Keep in mind that these settings can be read and followed only if crawlers are allowed to access the pages that include these settings.” (bản dịch) «Hãy nhớ rằng đó những settings có thể là đọc và followed chỉ nếu các crawler là được phép để access đó các trang đó bao gồm những settings.» Nhảy đến trích dẫn
- “If a page is disallowed from crawling through the robots.txt file, then any information about indexing or serving rules will not be found and will therefore be ignored.” (bản dịch) «Nếu một trang là disallowed từ crawling qua đó robots.txt file, thì bất kỳ information về lập chỉ mục hoặc serving rules sẽ không là được tìm thấy và sẽ do đó là đã bỏ qua.» Nhảy đến trích dẫn
- “For the
noindexrule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler.” (bản dịch) «Cho đónoindexrule để là effective, đó trang hoặc tài nguyên không được là blocked by một robots.txt file, và điều này có để là nếu không accessible để đó crawler.» — Google, “Block Search Indexing with noindex.” (bản dịch) «Block Tìm kiếm Lập chỉ mục với noindex.» Nhảy đến trích dẫn
Google — directives (verbatim các mô tả)
- “Do not show this page, media, or resource in search results.” (bản dịch) «Không cho thấy này trang, media, hoặc tài nguyên trong kết quả tìm kiếm.» —
noindex. Nhảy đến trích dẫn - “Equivalent to
noindex, nofollow.” (bản dịch) «Tương đương đểnoindex, nofollow.» —none. Nhảy đến trích dẫn - “Do not show a text snippet or video preview in the search results for this page.” (bản dịch) «Không cho thấy một text snippet hoặc video preview trong đó kết quả tìm kiếm cho này trang.» —
nosnippet. Nhảy đến trích dẫn - “Google is allowed to index the content of a page if it’s embedded in another page through iframes or similar HTML tags, in spite of a
noindexrule.” (bản dịch) «Google là được phép để chỉ mục đó nội dung of một trang nếu đây là embedded trong một sản phẩm khác trang qua iframes hoặc similar HTML tags, trong spite of mộtnoindexrule.» —indexifembedded. Nhảy đến trích dẫn - “A larger image preview, up to the width of the viewport, may be shown.” (bản dịch) «MỘT lớn hơn image preview, lên để đó width of đó viewport, có thể là shown.» —
max-image-preview:large. Nhảy đến trích dẫn - “Do not index images on this page.” (bản dịch) «Không chỉ mục images on này trang.» —
noimageindex. Nhảy đến trích dẫn - “You can designate textual parts of an HTML page not to be used as a snippet.” (bản dịch) «Bạn có thể designate textual parts of an HTML trang không để là dùng as một snippet.» —
data-nosnippet. Nhảy đến trích dẫn
Google — combining và resolving conflicts
- “You can create a multi-rule instruction by combining robots
metatag rules with commas or by using multiplemetatags.” (bản dịch) «Bạn có thể tạo một multi-rule instruction by combining robotsmetatag rules với commas hoặc by dùng multiplemetatags.» Nhảy đến trích dẫn - “In the case of conflicting robots rules, the more restrictive rule applies. For example, if a page has both
max-snippet:50andnosnippetrules, thenosnippetrule will apply.” (bản dịch) «Trong đó case of conflicting robots rules, đó hơn restrictive rule áp dụng. Ví dụ, nếu một trang có cả haimax-snippet:50vànosnippetrules, đónosnippetrule sẽ apply.» Nhảy đến trích dẫn - “For situations where multiple crawlers are specified along with different rules, the search engine will use the sum of the negative rules.” (bản dịch) «Cho situations nơi multiple các crawler là specified along với khác nhau rules, đó công cụ tìm kiếm sẽ dùng đó sum of đó negative rules.» Nhảy đến trích dẫn
Google — placement, crawler tokens, và case sensitivity
- “Google Search doesn’t enforce placement of meta robots in the HTML head and will respect robots meta tags in the body section of an HTML document as well.” (bản dịch) «Google Search không enforce placement of meta robots trong đó HTML head và sẽ respect robots meta tags trong đó thân phản hồi section of an HTML document as well.» Nhảy đến trích dẫn
- “Google supports two user agent tokens in the robots
metatag; other values are ignored.” (bản dịch) «Google hỗ trợ hai người dùng agent tokens trong đó robotsmetatag; other các giá trị là đã bỏ qua.» Nhảy đến trích dẫn - “Both the
nameand thecontentattributes are case-insensitive.” (bản dịch) «Cả hai đónamevà đócontentcác thuộc tính là case-insensitive.» Nhảy đến trích dẫn
Google — noarchive và khác lịch sử directives
- “The
noarchiverule is no longer used by Google Search to control whether a cached link is shown in search results, as the cached link feature no longer exists.” (bản dịch) «Đónoarchiverule là không lâu hơn dùng by Google Search để control liệu một được lưu đệm link là shown trong kết quả tìm kiếm, as đó được lưu đệm link feature không lâu hơn tồn tại.» Nhảy đến trích dẫn - “The
nocacherule isn’t used by Google Search.” (bản dịch) «Đónocacherule không dùng by Google Search.» Nhảy đến trích dẫn
Google — nosnippet và max-snippet reach AI Overviews và AI Chế độ
- “[nosnippet] applies to all forms of search results (at Google: web search, Google Images, Discover, AI Overviews, AI Mode) and will also prevent the content from being used as a direct input for AI Overviews and AI Mode.” (bản dịch) «[nosnippet] áp dụng để all forms of kết quả tìm kiếm (tại Google: web tìm kiếm, Google Images, Discover, AI Overviews, AI Chế độ) và sẽ cũng ngăn đó nội dung từ đang dùng as một trực tiếp input cho AI Overviews và AI Chế độ.» Nhảy đến trích dẫn
- “[max-snippet] applies to all forms of search results (such as Google web search, Google Images, Discover, Assistant, AI Overviews, AI Mode) and will also limit how much of the content may be used as a direct input for AI Overviews and AI Mode.” (bản dịch) «[max-snippet] áp dụng để all forms of kết quả tìm kiếm (such as Google web tìm kiếm, Google Images, Discover, Assistant, AI Overviews, AI Chế độ) và sẽ cũng limit cách nhiều of đó nội dung có thể là dùng as một trực tiếp input cho AI Overviews và AI Chế độ.» Nhảy đến trích dẫn
Google — không fixed timeframe cho noindex để take effect
- “If a page is still appearing in results, it’s probably because we haven’t crawled the page since you added the
noindexrule. Depending on the importance of the page on the internet, it may take months for Googlebot to revisit a page.” (bản dịch) «Nếu một trang là vẫn appearing trong kết quả, đây là probably vì we haven’t được crawl đó trang since bạn đã thêm đónoindexrule. Depending on đó importance of đó trang on đó internet, điều này có thể take months cho Googlebot để revisit một trang.» — Google, “Block Search Indexing with noindex.” (bản dịch) «Block Tìm kiếm Lập chỉ mục với noindex.» Nhảy đến trích dẫn
Google — X-Robots-Tag
- “The
X-Robots-Tagcan be used as an element of the HTTP header response for a given URL. Any rule that can be used in a robotsmetatag can also be specified as anX-Robots-Tag.” (bản dịch) «ĐóX-Robots-Tagcó thể là dùng as an element of đó HTTP header phản hồi cho một được cho URL. Bất kỳ rule đó có thể là dùng trong một robotsmetatag có thể cũng là specified as anX-Robots-Tag.» Nhảy đến trích dẫn - “You can use the
X-Robots-Tagfor non-HTML files like image files where the usage of robotsmetatags in HTML is not possible.” (bản dịch) «Bạn có thể dùng đóX-Robots-Tagcho non-HTML files như image files nơi đó usage of robotsmetatags trong HTML không phải có thể.» Nhảy đến trích dẫn
Patrick Stox — robots.txt là cho crawl control chỉ
- “Google was clear they want robots.txt for crawl control only. The biggest downside will probably be all the people who accidentally take their entire site out of the index.” (bản dịch) «Google đã là clear they muốn robots.txt cho crawl control chỉ. Đó biggest downside sẽ probably là all đó mọi người ai accidentally take của họ entire site out of đó chỉ mục.»
— Patrick Stox, on liệu Google nên thêm
noindexđể robots.txt, trong Search Engine Land. Nhảy đến trích dẫn
Patrick Stox — blocking xếp hạng các trang với robots.txt ( crawl-so với-chỉ mục proof)
- “Accidentally blocking pages (that Google already ranks) from being crawled using robots.txt probably isn’t going to have much impact on your rankings, and they will likely still show in the search results.” (bản dịch) «Accidentally blocking các trang (đó Google đã ranks) từ đang được crawl dùng robots.txt probably không going để có nhiều impact on của bạn thứ hạng, và they sẽ có khả năng vẫn cho thấy trong đó kết quả tìm kiếm.» Nhảy đến trích dẫn
- “We lost a position here or there and all of the featured snippets for the pages.” (bản dịch) «We lost một position ở đây hoặc ở đó và all of đó featured snippets cho đó các trang.» Nhảy đến trích dẫn
mỗi robots meta directive — bảng tra nhanh
đầy đủ đặt của Google-supported content các giá trị, Điều gì mỗi làm, và default.
| Directive | Điều gì nó làm | Default? |
|---|---|---|
all | Không restrictions on lập chỉ mục hoặc serving — implicit default | Có (Khi không tag) |
index | Cho phép trang trong kết quả tìm kiếm ( default; rarely được viết) | Implicit |
noindex | giữ điều này trang/media/tài nguyên out của kết quả tìm kiếm | Không |
follow | Follow links on điều này trang ( default; rarely được viết) | Implicit |
nofollow | không follow bất kỳ links on điều này trang (trang-wide) | Không |
none | Shorthand cho noindex, nofollow (không supported by Bing) | Không |
nosnippet | không hiển thị text snippet hoặc video preview; cũng chặn trang as trực tiếp input cho AI Overviews/AI Chế độ | Không |
max-snippet:[n] | Cap text snippet tại [n] characters (0 = none, -1 = không limit); cũng caps Cách nhiều có thể feed AI Overviews/AI Chế độ trực tiếp | Không |
max-image-preview:[setting] | Cap image-preview size: none / standard / large | Không |
max-video-preview:[n] | Cap video preview tại [n] seconds (0 = none, -1 = không limit) | Không |
notranslate | không offer translation của điều này trang trong kết quả | Không |
noimageindex | không chỉ mục images on điều này trang | Không |
unavailable_after:[date/time] | Drop trang từ kết quả sau khi được cho date/time | Không |
indexifembedded | Cho phép lập chỉ mục của nội dung embedded qua iframe ngay cả với noindex | Không |
Lịch sử — Google không lâu hơn dùng những điều này:
| Directive | Status |
|---|---|
noarchive | Không lâu hơn được sử dụng — được lưu đệm-link feature nó controlled không lâu hơn tồn tại |
nocache | không được sử dụng by Google Search (some engines treated nó as noarchive synonym) |
nositelinkssearchbox | Không lâu hơn được sử dụng — sitelinks tìm kiếm box feature nó controlled không lâu hơn tồn tại |
Paragraph-cấp độ ( HTML thuộc tính, không content giá trị):
| Thuộc tính | Điều gì nó làm | nơi |
|---|---|---|
data-nosnippet | giữ cụ thể passage out của snippet | On span, div, section |
** syntax**
<!-- one tag, comma-separated -->
<meta name="robots" content="noindex, nofollow">
<!-- target one engine -->
<meta name="googlebot" content="noindex">
<!-- the HTTP-header equivalent, for non-HTML / at scale -->
X-Robots-Tag: noindex
X-Robots-Tag: googlebot: noindex, nofollowFast facts
- Không tag tại all → default
index, follow. none=noindex, nofollow— nhưng Bing không hỗ trợnone; ghi nó out.max-*,noimageindex,notranslate,indexifembedded,unavailable_afterlà effectively Google-chỉ.- tag là HTML-chỉ;
X-Robots-Tagheader carries giống nhau rules cho PDFs, images, và toàn bộ directories. - Google đọc chính xác hai crawler-name tokens —
googlebotvàgooglebot-news— và bỏ qua mọi thứ khác, including khác engines’ tokens nhưbingbot. - Head placement là portable convention, nhưng Google cụ thể cũng
respects robots meta tag placed trong
<body>. name/contentvàX-Robots-Tagcác giá trị là case-insensitive để Google.
mental models
1. Crawl-sau đó-obey — tag chỉ hoạt động nếu trang là fetchable.
Google có để crawl trang để đọc tag. Bất cứ điều gì đó chặn fetch
(robots.txt disallow, auth, máy chủ lỗi) có nghĩ là tag là không bao giờ seen. So
trước khi bạn trust noindex, xác nhận URL là crawlable. corollary: không bao giờ
Disallow URL bạn’re trying để noindex.
2. Three tools, three jobs — không mix them lên.
robots.txt= crawl control (liệu bots fetch trang).- Robots meta tag / X-Robots-Tag = chỉ mục & phục vụ control (Điều gì happens sau khi nó fetched).
- Authentication = secrecy (
noindextrang là vẫn công khai). Match job để tool. phần lớn phổ biến thất bại là sử dụngrobots.txtđể try để deindex — đó meta tag job.
3. decision rule cho removing trang.
Muốn nó out của tìm kiếm? Leave nó crawlable và thêm noindex — và
không cũng block nó trong robots.txt, và không cũng canonical nó để khác
URL. Muốn bots để skip space hoàn toàn (và bạn không care về lập chỉ mục)?
robots.txt disallow. hai không phải interchangeable.
4. Conflict resolution — phần lớn restrictive wins.
Khi rules collide, stricter một áp dụng (nosnippet beats max-snippet:50).
không try để out-clever điều này với combinations; assume tightest rule là một
đó takes effect.
5. googlebot so với robots — sum của negatives, không override.
cho overlap, Googlebot adds lên restrictions từ generic robots
tag và googlebot-named tag thay vì picking một. robots: nofollow +
googlebot: noindex ⇒ Googlebot nhận noindex, nofollow. ( crawler-named tag
làm take precedence over generic tag nơi họ đặt giống nhau directive
differently — nhưng union là rule để remember.)
6. HTML trang → meta tag; mọi thứ khác → header.
nếu nó có <head>, sử dụng <meta> tag. nếu nó PDF, image, bất kỳ non-HTML
file, hoặc bạn cần cover toàn bộ directory, sử dụng X-Robots-Tag header —
giống nhau rule vocabulary, khác phân phối.
Nhật ký thay đổi
Đã cập nhật 8 thg 8, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.
Đã cập nhật 18 thg 7, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.