Hướng dẫn về X-Robots-Tag

Đó X-Robots-Tag HTTP header carries đó giống nhau directives as đó robots meta tag, so bạn có thể noindex PDFs, images, và videos. Apache, Nginx, và Cloudflare configs plus cách verify với curl.

Xuất bản lần đầu: 23 thg 6, 2026 · Cập nhật lần cuối: 8 thg 8, 2026 · Advanced
Ngôn ngữ
1 tín hiệu bằng chứng trên trang này

Đó X-Robots-Tag là đó HTTP-header twin of đó robots meta tag: giống nhau directive vocabulary (noindex, nofollow, nosnippet, và đó rest), nhưng delivered trong đó header phản hồi so điều này hoạt động on files với không HTML <head> — PDFs, images, videos, bất cứ điều gì. Bạn set điều này tại đó máy chủ (Apache, Nginx, hoặc một CDN/Worker), so một rule có thể noindex mỗi PDF on đó site. Đó single hầu hết phổ biến mistake là cũng blocking đó file trong robots.txt — mà hides đó header, so Google không bao giờ crawl đó file để đọc đó noindex và đó directive là đã bỏ qua. Verify đó trực tiếp header với curl -I, và watch cho một CDN serving một khác nhau header hơn của bạn origin (đó 'Noindex detected trong X-Robots-Tag' Search Console failure).

TL;DR — Đó X-Robots-Tag là đó HTTP-header phân phối of đó giống nhau directive vocabulary as đó robots meta tag — “Any rule that can be used in a robots meta tag can also be specified as an X-Robots-Tag.” (bản dịch) «Bất kỳ rule đó có thể là dùng trong một robots tag có thể cũng là specified as an .» Điều này tồn tại cho một big reason: non-HTML các tài nguyên (PDFs, images, videos) có không <head>, so đó header là đó supported way để noindex them. Bạn set điều này máy chủ-side (Apache <FilesMatch>, Nginx add_header, hoặc một Cloudflare Transform Rule / Worker), so một rule covers một toàn bộ file loại. Tùy chọn đích một crawler với X-Robots-Tag: googlebot: noindex; với không người dùng-agent điều này áp dụng để all of them. Đó spine caveat, shared với đó meta tag: đó file phải được crawlable — một robots.txt block hides đó header, so đó noindex là “not found and… therefore ignored.” (bản dịch) «không tìm thấy và… do đó đã bỏ qua.» Verify với curl -I, và watch cho CDN-so với-origin header drift.

Evidence for this claim Google supports robots directives in the X-Robots-Tag HTTP response header, including for non-HTML resources such as PDFs. Scope: Google crawling and indexing controls delivered by HTTP header. Confidence: high · Verified: Google Search Central: X-Robots-Tag Evidence for this claim Google must be allowed to crawl a resource to discover and apply its X-Robots-Tag noindex rule. Scope: A robots.txt block can prevent Google from seeing the response header. Confidence: high · Verified: Google Search Central: Combining crawling and indexing rules

Điều gì đó X-Robots-Tag thực ra là

Đó X-Robots-Tag là đó HTTP header phản hồi version of đó robots <meta> tag. đó là đó toàn bộ concept. Google là rõ ràng đó hai là interchangeable trong vocabulary: “The X-Robots-Tag can be used as an element of the HTTP header response for a given URL. Any rule that can be used in a robots meta tag can also be specified as an X-Robots-Tag.” (bản dịch) «Đó có thể là dùng as an element of đó HTTP header phản hồi cho một được cho URL. Bất kỳ rule đó có thể là dùng trong một robots tag có thể cũng là specified as an .» MỘT phản hồi carrying điều này looks như này:

Evidence for this claim For Google, every robots-meta rule that Google supports can also be supplied in the `X-Robots-Tag` response header; this is a Google contract, not proof of identical vocabulary across all crawlers. Scope: production HTML and HTTP responses Confidence: high · Verified: Robots meta tag, data-nosnippet, and X-Robots-Tag specifications
HTTP/1.1 200 OK
Date: Tue, 25 May 2010 21:42:43 GMT
(…)
X-Robots-Tag: noindex
(…)

So nếu bạn đã understand đó robots meta tag, bạn đã understand đó directives ở đây — noindex, nofollow, nosnippet, và đó rest behave identically cho Google. Điều duy nhất đó thay đổi là đó phân phối mechanism: header, không HTML.

Đó sự tương đương là một Google-được ghi lại contract, không một universal web tiêu chuẩn — đó X-Robots-Tag là một de facto header với không IETF đặc tả backing điều này. Google says so trực tiếp on đó giống nhau doc: “It is possible that these rules may not be treated the same by all other search engines.” (bản dịch) «Điều này là có thể đó những rules có thể không là treated đó giống nhau by all other các công cụ tìm kiếm.» Treat “same directives as the meta tag” (bản dịch) «giống nhau directives as đó meta tag» as đúng cho Google (và, theo đó Other Engines section dưới, largely đúng cho Bing cốt lõi set) — không as một bảo đảm cho mỗi crawler, including AI các crawler, đó có thể đọc này header.

Vì sao điều này tồn tại — non-HTML files

Đó defining dùng case là files đó có không HTML <head> để drop một meta tag vào. Google hướng dẫn là trực tiếp: “To block indexing of non-HTML resources, such as PDF files, video files, or image files, use the X-Robots-Tag response header instead.” (bản dịch) «Để block lập chỉ mục of non-HTML các tài nguyên, such as PDF files, video files, hoặc image files, dùng đó header phản hồi thay vì.» đó là đó headline. MỘT PDF, an image, hoặc một video không thể carry một robots meta tag — có nowhere để put điều này — so đó header là đó supported route để noindex them.

có một second, quieter reason: đây là set máy chủ-side, so một rule có thể apply một directive để an entire class of files (mỗi .pdf on đó site) hoặc site-wide, không có bạn editing mỗi tài nguyên. đó là powerful, và — as we’ll see — cũng nơi đó foot-guns trực tiếp.

Worked ví dụ — noindexing một single PDF: chẳng hạn https://example.com/whitepapers/old-report.pdf là cho thấy lên trong tìm kiếm và bạn muốn điều này đã biến mất. Hai conditions cả hai có để hold: (1) đó tài nguyên phải stay crawlable — không Disallow on điều này trong robots.txt — và (2) đó máy chủ phản hồi cho đó chính xác URL phải carry X-Robots-Tag: noindex. Set đó rule (Apache/Nginx/ Cloudflare các ví dụ dưới), xác nhận với curl -I đó trực tiếp phản hồi thực ra trả về đó header (see đó verification section), thì chờ cho Google để recrawl đó URL. Skip either condition và đó PDF vẫn giữ được lập chỉ mục: skip crawlability và Google không bao giờ đọc đó header; skip đó header và có không có gì telling Google để drop điều này.

Không turn “non-HTML” vào “wrong file type.” (bản dịch) «sai file loại.» Google publishes một rộng list of indexable file types, including PDF, phổ biến office formats, text, XML, và several nguồn-code formats. MỘT crawler nên bảo toàn three tách biệt observations: đó URL extension, đó declared Content-Type, và đó bytes thực ra đã trả về. MỘT mismatch là một phân phối warning; một legitimate supported document không phải an HTML-trang failure và nên có HTML-chỉ kiểm tra marked không applicable. Dùng X-Robots-Tag hoặc an HTTP Link canonical khi những controls là needed cho đó non-HTML tài nguyên.

Size limits cũng cần an engine-và-basis label. Google March 2026 crawler cập nhật documents một 2 MB theo-URL cutoff cho Googlebot (Tìm kiếm), một 64 MB PDF cutoff, và một 15 MB default cho unspecified Google crawler clients. Những là decimal MB tài nguyên limits, không interchangeable với một tool decoded-thân phản hồi MiB safety cap. See Google hiện tại crawler limits trước chuyển thành một sản phẩm limit vào an SEO diagnosis.

Meta tag so với X-Robots-Tag — khi nào nên dùng mà

Đó decision là đơn giản khi bạn frame điều này khoảng đó <head>:

  • HTML trang bạn control đó markup of → dùng đó robots meta tag. đây là easier để set, easier để đọc, và theo-trang.
  • Non-HTML file (PDF, image, video, feed) → dùng đó X-Robots-Tag header. có không other option; những có không <head>.
  • Bạn’d rather control điều này máy chủ-side / site-wide, hoặc bạn không thể easily touch mỗi trang HTML (CDN, app layer) → đó header là đó cleaner lever ngay cả cho HTML.

họ là không rivals và một không “stronger” — giống nhau directives, khác nhau reach.

Đó directives bạn có thể dùng

Đó X-Robots-Tag accepts đó giống nhau directive vocabulary as đó meta tag. Đó đầy đủ supported set:

  • noindex — không chỉ mục này tài nguyên.
  • nofollow — không follow links từ điều này.
  • none — shorthand cho noindex, nofollow.
  • all — không restrictions (đó default).
  • nosnippet — không text snippet hoặc video preview trong kết quả (đó trang-cấp độ cousin of data-nosnippet, mà scopes điều này để part of một trang).
  • max-snippet: [number] — cap đó snippet length.
  • max-image-preview: [none | standard | large] — cap đó preview image size.
  • max-video-preview: [number] — cap đó video preview length trong seconds.
  • noarchive — không được lưu đệm link.
  • noimageindex — không chỉ mục images on đó trang.
  • notranslate — không offer translation trong kết quả.
  • indexifembeddedkhông chỉ mục này tài nguyên on của nó own, nhưng làm cho phép lập chỉ mục nơi đây là embedded (e.g. trong an iframe). Pair điều này với noindex.
  • unavailable_after: [date/time] — drop đó URL từ kết quả sau một date. Đó date phải được trong một widely adopted format (RFC 822, RFC 850, hoặc ISO 8601), e.g.:
X-Robots-Tag: unavailable_after: 25 Jun 2010 15:00:00 PST

Presenting đó bảng khi là deliberate: đây là đó giống nhau list shared by đó meta tag, mà là đó sạch nhất way để think về đó mối quan hệ giữa đó hai.

Targeting một cụ thể crawler

Bạn có thể tùy chọn prefix đó directive với một người dùng-agent token: Google notes đó header “may optionally specify a user agent before the rules,” (bản dịch) «có thể tùy chọn specify một người dùng agent trước đó rules,» và đó “Rules specified without a user agent are valid for all crawlers. The HTTP header, the user agent name, and the specified values are not case sensitive.” (bản dịch) «Rules specified không có một người dùng agent là hợp lệ cho all các crawler. Đó HTTP header, người dùng agent name, và đó specified các giá trị không phải case sensitive.» So này là hợp lệ syntax để noindex chỉ cho Google:

Evidence for this claim Google's header syntax permits an optional crawler name before a rule list, and Google treats header names, crawler names and directive values as case-insensitive. Scope: production HTML and HTTP responses Confidence: high · Verified: Robots meta tag, data-nosnippet, and X-Robots-Tag specifications
X-Robots-Tag: googlebot: noindex

Và bạn có thể stack lines cho khác nhau bots. (Một note on attribution: Google own được ghi lại ví dụ of theo-crawler targeting dùng X-Robots-Tag: googlebot: nofollow paired với X-Robots-Tag: otherbot: noindex, nofollow — so googlebot: noindex là correct, supported syntax, chỉ không đó verbatim line trong Google sample.)

Multiple rules và conflicting directives

bạn là không limited để một directive theo phản hồi. Google documents này trực tiếp: “Multiple rules may be combined in a comma-separated list or in separate meta tags. These rules are case-insensitive.” (bản dịch) «Multiple rules có thể là combined trong một comma-separated list hoặc trong tách biệt meta tags. Những rules là case-insensitive.» — đó giống nhau là đúng cho đó header: gửi X-Robots-Tag: noindex, nofollow as một comma-separated line, hoặc repeat đó header trường (X-Robots-Tag: noindex on một line, X-Robots-Tag: nofollow on một sản phẩm khác). Cả hai là hợp lệ; đó máy chủ config các ví dụ dưới dùng đó comma-separated form.

Khi hai rules sẽ conflict, Google resolves điều này trong favor of đó tighter restriction: “In the case of conflicting robots rules, the more restrictive rule applies. For example, if a page has both max-snippet:50 and nosnippet rules, the nosnippet rule will apply.” (bản dịch) «Trong đó case of conflicting robots rules, đó hơn restrictive rule áp dụng. Ví dụ, nếu một trang có cả hai và rules, đó rule sẽ apply.» đó là Google được ghi lại behavior — đó giống nhau “not treated the same by all other search engines” (bản dịch) «không treated đó giống nhau by all other các công cụ tìm kiếm» caveat trên áp dụng để conflict resolution cũng, so không assume mỗi crawler picks đó stricter rule cùng cách.

Evidence for this claim Google accepts multiple X-Robots-Tag rules as a comma-separated list or repeated header fields, and resolves conflicting rules by applying the more restrictive one; this behavior is documented for Google specifically, not guaranteed across all crawlers. Scope: Google's handling of multiple/conflicting robots meta and X-Robots-Tag directives. Confidence: high · Verified: Google Search Central: repeated and comma-separated X-Robots-Tag rules Supports: Multiple X-Robots-Tag rules may use repeated headers or a comma-separated list. Google Search Central: conflicting robots rules Supports: Google applies the more restrictive rule when robots rules conflict. Google Search Central: cross-engine treatment may differ Supports: Other search engines may treat these rules differently.

Cách set điều này (máy chủ config)

Những là known-good, battle-tested patterns I’d reach cho — treat mỗi snippet dưới as an environment-cụ thể ví dụ, không một portable recipe: inheritance rules, route matching, các chuyển hướng, phản hồi các mã trạng thái, và bất kỳ CDN hoặc application layer trong front of đó máy chủ có thể all thay đổi điều gì thực ra reaches đó crawler. Xác nhận đó delivered header on của bạn own stack trước shipping, không chỉ đó config file. Configs và snippets là trong đó Scripts tab, sao chép và dán ready, cho Apache (<FilesMatch> + mod_headers), Nginx (location

  • add_header), và Cloudflare (một Transform Rule hoặc một Worker). Phiên bản ngắn gọn: một rule, khớp on đó file extension, sets X-Robots-Tag: noindex on mỗi file of đó loại. Để đích một single crawler, người dùng-agent token goes bên trong đó header giá trị (Header set X-Robots-Tag "googlebot: noindex"), không as một tách biệt mechanism.

Đó #1 mistake — không block đó file trong robots.txt

Này là đó spine of đó toàn bộ topic, so I’ll là loud về điều này. Mọi người ai muốn một PDF đã biến mất thường làm hai điều tại khi: block điều này trong robots.txt set X-Robots-Tag: noindex. Đó defeats itself. Google có để crawl đó file để đọc đó header — và đó robots.txt block dừng đó crawl.

Google own wording: “robots meta tags and X-Robots-Tag HTTP headers are discovered when a URL is crawled. If a page is disallowed from crawling through the robots.txt file, then any information about indexing or serving rules will not be found and will therefore be ignored.” (bản dịch) «robots meta tags và X-Robots-Tag HTTP các header là discovered khi một URL là được crawl. Nếu một trang là disallowed từ crawling qua đó robots.txt file, thì bất kỳ information về lập chỉ mục hoặc serving rules sẽ không là được tìm thấy và sẽ do đó là đã bỏ qua.» Và đó corollary: “If indexing or serving rules must be followed, the URLs containing those rules cannot be disallowed from crawling.” (bản dịch) «Nếu lập chỉ mục hoặc serving rules phải được followed, đó URLs containing những rules không thể là disallowed từ crawling.»

Restated cho noindex cụ thể: “For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler. If the page is blocked by a robots.txt file or the crawler can’t access the page, the crawler will never see the noindex rule, and the page can still appear in search results.” (bản dịch) «Cho đó rule để là effective, đó trang hoặc tài nguyên không được là blocked by một robots.txt file, và điều này có để là nếu không accessible để đó crawler. Nếu đó trang là blocked by một robots.txt file hoặc đó crawler không thể access đó trang, đó crawler sẽ không bao giờ see đó rule, và đó trang có thể vẫn xuất hiện trong kết quả tìm kiếm.» Này là đó giống nhau crawling-so với-lập chỉ mục phân biệt đó chạy qua robots.txt và đó meta tag: robots.txt controls crawling; đó X-Robots-Tag controls lập chỉ mục; và một crawl block hides đó lập chỉ mục rule. Để deindex một PDF: cho phép crawling, serve X-Robots-Tag: noindex, và chờ cho đó re-crawl.

(Note đó flip side cũng: một robots.txt Disallow on của nó own không deindex một file either — một blocked PDF có thể vẫn là được lập chỉ mục nếu other các trang link để điều này, chỉ không có một snippet. Blocking và noindexing solve khác nhau các vấn đề.)

Cách verify điều này với curl

không trust đó một config “nên” hoạt động — kiểm tra đó trực tiếp header. curl -I (hoặc curl -sI … | grep -i x-robots-tag) prints đó các header phản hồi không có downloading đó thân phản hồi:

curl -sI https://example.com/file.pdf | grep -i x-robots-tag

Bạn có thể cũng spoof đó bot view để see điều gì Googlebot sẽ nhận (curl -A "Googlebot" -sI …). Hầu hết các hướng dẫn point bạn tại một trình duyệt extension hoặc Screaming Frog cho này; curl là nhanh hơn và scriptable, và đầy đủ commands là trong đó Scripts tab.

Một caveat on -I: điều này gửi một HEAD yêu cầu, và một máy chủ hoặc CDN rule đó chỉ fires on GET sẽ không cho thấy lên trong một HEAD phản hồi. Treat curl -I as một fast đầu tiên kiểm tra, không đó cuối word — xác nhận đó giống nhau header on một real curl GET yêu cầu (curl -sD - -o /dev/null https://example.com/file.pdf | grep -i x-robots-tag), và kiểm tra điều này sau các chuyển hướng và on đó công khai CDN-fronted URL, không chỉ origin.

Watch cho CDN-so với-origin drift

Vì đó header là set máy chủ- và CDN-side, đó giá trị Googlebot nhận có thể differ từ điều gì của bạn origin emits — an edge rule có thể thêm, strip, hoặc bộ nhớ đệm một header. Đó thực tế chế độ lỗi ở đây là đó “Noindex detected in X-Robots-Tag” (bản dịch) «Noindex detected trong X-Robots-Tag» message trong Search Console: một CDN hoặc edge config serving an X-Robots-Tag: noindex — ngay cả briefly hoặc by mistake — nhận logged by Google as một noindex tín hiệu và có thể deindex các trang bạn không bao giờ meant để touch. Đó lesson: luôn curl -I đó công khai URL behind đó CDN, không chỉ origin, và treat edge rules với đó giống nhau care as origin config.

Other engines

Bing hỗ trợ đó X-Robots-Tag header và broadly đó giống nhau cốt lõi directives — noindex, nofollow, noarchive/nocache (Bing dùng nocache as một synonym cho noarchive), nosnippet, và đó preview-limit directives. Nếu bạn cần Bing chính xác supported list, xác nhận điều này on Bing trực tiếp help trang — của họ tài liệu là JavaScript-được kết xuất và worth eyeballing trực tiếp trước relying on một cụ thể directive.

Nơi này sits

Đó X-Robots-Tag là đó non-HTML half of on-trang lập chỉ mục control: đó robots meta tag cho các trang, này header cho mọi thứ khác. Điều này leans on đó giống nhau noindex, nosnippet / max-snippet / max-image-preview vocabulary, và điều này phụ thuộc vào crawling đang được phép — mà ties điều này straight lại để robots.txt và đó rộng hơn lập chỉ mục story. Nếu bạn nghĩ ra ở đây từ đó crawling side, đó takeaway là đó caveat; nếu bạn nghĩ ra từ đó meta-tag side, đó takeaway là đó giống nhau directives reach PDFs và media qua đó header.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.