Hướng dẫn về Noindex

Noindex giữ một trang out of kết quả tìm kiếm — nhưng chỉ nếu Google có thể crawl điều này. Đó hai hợp lệ các phương thức, đó robots.txt trap, và cách verify điều này worked.

Xuất bản lần đầu: 23 thg 6, 2026 · Cập nhật lần cuối: 8 thg 8, 2026 · Advanced
Ngôn ngữ

Noindex là đó directive đó giữ một trang out of đó chỉ mục, so điều này sẽ không xuất hiện trong kết quả tìm kiếm. Có hai hợp lệ ways để set điều này: đó robots meta tag (`<meta name="robots" content="noindex">`) và đó `X-Robots-Tag: noindex` HTTP header (đó chỉ option cho non-HTML files như PDFs). Đó single biggest mistake: một trang blocked trong robots.txt không thể là noindexed, vì Google không bao giờ crawl điều này để see đó rule — so để xóa một trang bạn có để cho phép crawling và serve noindex. không put noindex trong robots.txt (unsupported since Sept 1, 2019); nếu noindex và canonical coexist, treat đó as an intent kiểm tra thay vì an tự động lỗi; và remember deindexing chỉ happens sau một recrawl.

TL;DR — noindex xóa một trang từ đó chỉ mục qua một of hai hợp lệ các phương thức: đó robots meta tag (<meta name="robots" content="noindex">) hoặc đó X-Robots-Tag: noindex HTTP header (bắt buộc cho non-HTML files như PDFs). Đó load-bearing gotcha: một trang blocked trong robots.txt không thể là noindexed — Google không bao giờ crawl điều này để see đó rule, “the crawler will never see the noindex rule,” (bản dịch) «đó crawler sẽ không bao giờ see đó rule,» và một linked URL có thể stay được lập chỉ mục. So để xóa một trang, cho phép crawling và serve noindex. không put noindex trong robots.txt (unsupported since Sept 1, 2019), review noindex với một canonical pointing elsewhere as potentially conflicting, và know đó deindexing chỉ happens sau một recrawl — Google own hướng dẫn says một thấp-priority trang có thể take months. Theo một 2017 Mueller comment (không được ghi lại policy), dài-term noindex,follow tends để behave như noindex,nofollow khi đó trang drops từ đó chỉ mục. Verify trong GSC dưới “URL marked ‘noindex’.” (bản dịch) «URL được đánh dấu noindex.»

Điều gì noindex là — chỉ mục control, không crawl control

noindex là đó chính chỉ mục-control directive. Google own definition of đó rule là một line: “Do not show this page, media, or resource in search results.” (bản dịch) «Không cho thấy này trang, media, hoặc tài nguyên trong kết quả tìm kiếm.» Khi đây là honored, đó effect là total — “When Googlebot crawls that page and extracts the tag or header, Google will drop that page entirely from Google Search results, regardless of whether other sites link to it.” (bản dịch) «Khi Googlebot crawl đó trang và extracts đó tag hoặc header, Google sẽ drop đó trang hoàn toàn từ Google Search kết quả, regardless of liệu other các trang link để điều này.» Evidence for this claim Google's noindex rule prevents the page, media, or resource from appearing in Google Search results after Google sees the rule. Scope: Google Search; noindex is not an access-control or privacy mechanism. Confidence: high · Verified: Google Search Central: Block Search indexing with noindex

Giữ một phân biệt front of mind, vì gần như mỗi noindex mistake xuất hiện từ blurring điều này: noindex controls lập chỉ mục; robots.txt controls crawling. họ là khác nhau stages of đó pipeline. I put điều này way trong my Ahrefs hướng dẫn on removing URLs: “Crawling is not the same thing as indexing. Even if Google is blocked from crawling pages, if there are any internal or external links to a page they can still index it.” (bản dịch) «Crawling không phải cùng một điều as lập chỉ mục. Ngay cả khi Google là blocked từ crawling các trang, nếu có bất kỳ internal hoặc liên kết bên ngoài để một trang they có thể vẫn chỉ mục điều này.» Đó sentence là đó toàn bộ reason đó rest of này bài viết tồn tại.

Microsoft cho giống nhau directive additional Bing-cụ thể consequence: nội dung marked noindex là cũng excluded từ Microsoft foundation-model training. prerequisite vẫn matters—Bingbot phải là được phép để crawl và xử lý trang-cấp độ directive. robots.txt block plus noindex là làm đó không proof đó either deindexing hoặc training opt-out có là applied.

Evidence for this claim Microsoft says content marked noindex is not included in the Bing index and is not used to train its generative AI foundation models. Scope: Bing and Microsoft foundation-model use; Bingbot must be able to crawl and process the directive before the outcome can be inferred. Confidence: high · Verified: Bing Webmaster Blog: New controls for Bing Chat

hai hợp lệ phân phối các phương thức

có chính xác hai, và noindex trong robots.txt là không một của them (nhiều hơn on đó dưới).

Phương thức 1 — robots meta tag. cho HTML trang, place điều này trong <head>:

<meta name="robots" content="noindex">

Google instruction là verbatim: “To prevent all search engines that support the noindex rule from indexing a page on your site, place the following <meta> tag into the <head> section of your page.” (bản dịch) «Để ngăn all các công cụ tìm kiếm đó hỗ trợ đó rule từ lập chỉ mục một trang trên trang web của bạn, place đó sau thẻ meta tag vào đó thẻ head section of trang của bạn.» Đó robots giá trị targets all các crawler đó hỗ trợ đó rule; đổi trong googlebot để đích chỉ Google (<meta name="googlebot" content="noindex">).

Phương thức 2 — X-Robots-Tag HTTP header. giống nhau directive, được gửi trong phản hồi header thay vì markup:

X-Robots-Tag: noindex

Này là đó chỉ way để noindex non-HTML các tài nguyên, vì có không <head> để host một meta tag. Google: “A response header can be used for non-HTML resources, such as PDFs, video files, and image files.” (bản dịch) «MỘT header phản hồi có thể là dùng cho non-HTML các tài nguyên, such as PDFs, video files, và image files.» Và từ đó robots spec: bạn có thể dùng đó X-Robots-Tag “for non-HTML files like image files where the usage of robots meta tags in HTML is not possible.” (bản dịch) «cho non-HTML files như image files nơi đó usage of robots tags trong HTML không phải có thể.» Evidence for this claim Google supports noindex in an HTML robots meta tag or an X-Robots-Tag HTTP response header. Scope: Google Search delivery methods; the HTTP header is applicable to non-HTML resources as well as HTML. Confidence: high · Verified: Google Search Central: Robots meta tag and X-Robots-Tag specifications

Một placement note: put đó meta tag trong đó <head> — đó là đó tiêu chuẩn, safest spot và điều gì Google cách-để cho thấy. Google spec trang làm chẳng hạn điều này “doesn’t enforce placement of meta robots in the HTML head and will respect robots meta tags in the body section of an HTML document as well,” (bản dịch) «không enforce placement of meta robots trong đó HTML head và sẽ respect robots meta tags trong đó thân phản hồi section of an HTML document as well,» nhưng không rely on đó as của bạn chính phương thức; một stray <meta> tag some CMS injects vào đó <body> có thể noindex một trang by accident chỉ as easily as một bạn meant để thêm để đó <head>.

Since header là configured tại máy chủ cấp độ, nó varies by stack. Hai phổ biến các ví dụ cho noindexing mỗi PDF on trang web:

Apache (.htaccess hoặc vhost):

<FilesMatch "\.pdf$">
  Header set X-Robots-Tag "noindex"
</FilesMatch>

Nginx (server/location block):

location ~* \.pdf$ {
  add_header X-Robots-Tag "noindex";
}

#1 mistake — noindex + robots.txt block

Noindex is crawl-then-obey: keep the URL fetchable long enough for the directive to be processed.

The same page contains a meta robots noindex directive. With crawling allowed, Google can fetch the page, see noindex, and remove the URL after processing. With crawling blocked in robots.txt, Google cannot see noindex and the linked URL may remain in results.

Đây là thất bại chế độ I see phần lớn, so ở đây mechanism trong đầy đủ. noindex tag lives on trang; Google có để fetch trang để đọc nó. Google trạng thái requirement trực tiếp:

“For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler. If the page is blocked by a robots.txt file or the crawler can’t access the page, the crawler will never see the noindex rule, and the page can still appear in search results, for example if other pages link to it.” (bản dịch) «Cho đó rule để là effective, đó trang hoặc tài nguyên không được là blocked by một robots.txt file, và điều này có để là nếu không accessible để đó crawler. Nếu đó trang là blocked by một robots.txt file hoặc đó crawler không thể access đó trang, đó crawler sẽ không bao giờ see đó rule, và đó trang có thể vẫn xuất hiện trong kết quả tìm kiếm, ví dụ nếu other các trang link để điều này.»

Put ngay cả hơn bluntly: “We have to crawl your page in order to see <meta> tags and HTTP headers.” (bản dịch) «We có để crawl trang của bạn trong order để see thẻ meta tags và HTTP các header.» Không crawl, không rule.

robots.txt là đó hầu hết phổ biến way một trang ends lên uncrawlable, nhưng Google wording covers hơn ground hơn đó — điều này cũng says “the crawler can’t access the page,” (bản dịch) «đó crawler không thể access đó trang,» mà bao gồm repeated máy chủ các lỗi (5xx), timeouts, và an unintended authentication wall trong front of đó trang. Bất kỳ of những silently breaks noindex cùng cách một robots.txt block làm.

So đó instinct để “block it in robots.txt and noindex it, just to be safe” (bản dịch) «block điều này trong robots.txt noindex điều này, chỉ để là safe» là chính xác backwards — đó block ngăn đó crawl, đó crawl là điều gì reveals đó noindex, và đó trang có thể sit trong đó chỉ mục indefinitely (thường as một mô tả-ít hơn URL). Trong Google Search Console này cho thấy lên as đó “Indexed, though blocked by robots.txt” (bản dịch) «Được lập chỉ mục, though Bị chặn bởi robots.txt» status — một trang bạn blocked đó đã nhận được lập chỉ mục anyway vì điều gì đó links để điều này.

** khắc phục:** unblock trang trong robots.txt, giữ noindex on nó, và let Google recrawl. chỉ sau khi trang có dropped từ chỉ mục — nếu bạn sau đó muốn để save crawl hoàn toàn — là nó safe để thêm disallow.

Worked deployment ví dụ: staging trang web đó sẽ không disappear

redesign launches từ staging.example.com. staging templates đã contain noindex, nhưng deployment checklist cũng adds:

User-agent: *
Disallow: /

đó feels như hai layers của protection. nó là thực ra trap nếu Google đã discovered staging các URL qua shared QA link, old sitemap, công khai ticket, hoặc link trong copied production nội dung. disallow ngăn tiếp theo crawl, so Google không thể xác nhận noindex; hostname có thể linger as thin, URL-chỉ kết quả.

cleanup sequence là: xóa disallow, giữ noindex on mỗi staging phản hồi, xác nhận trực tiếp phản hồi là crawlable và exposes directive, yêu cầu recrawling cho representative sample, và monitor hostname cho đến khi nó drops out. sau đó put environment behind authentication. Authentication là durable quyền riêng tư control; noindex là chỉ tìm kiếm-chỉ mục control.

noindex so với nofollow so với disallow

Three directives mọi người constantly conflate. họ operate tại khác stages:

  • noindexchỉ mục control. Trang là được crawl, kept out of kết quả. Google definition: “Do not show this page, media, or resource in search results.” (bản dịch) «Không cho thấy này trang, media, hoặc tài nguyên trong kết quả tìm kiếm.»
  • nofollowlink control. Google: “Do not follow the links on this page.” (bản dịch) «Không follow đó links on này trang.» Điều này says không có gì về lập chỉ mục đó trang itself.
  • disallow (robots.txt) — crawl control. Dừng đó fetch hoàn toàn. Điều này là không an chỉ mục control — một disallowed URL có thể vẫn là được lập chỉ mục nếu đây là linked.

có cũng none, mà Google documents as “Equivalent to noindex, nofollow.” (bản dịch) «Tương đương để .» Và khi directives conflict, đó spec là clear: “In the case of conflicting robots rules, the more restrictive rule applies.” (bản dịch) «Trong đó case of conflicting robots rules, đó hơn restrictive rule áp dụng.» (Đầy đủ bảng on đó Các bảng tra nhanh tab.)

Treat noindex với rel=canonical as intent kiểm tra

Putting noindexrel="canonical" on đó cùng trang không phải tự động không hợp lệ. Điều này làm tạo một configuration worth reviewing: một canonical asks Google để consolidate các tín hiệu, trong khi noindex asks cho này URL để là excluded. Cho chọn giữa duplicates, dùng đó canonical tag — Google cụ thể advises so với dùng noindex cho điều này: “We don’t recommend using noindex to prevent selection of a canonical page within a single site, because it will completely block the page from Search.” (bản dịch) «We không khuyến nghị dùng để ngăn selection of một canonical trang trong một single site, vì điều này sẽ completely block đó trang từ Tìm kiếm.» Note đó phạm vi: Google caution là cụ thể về dùng noindex để pick mà duplicate wins as canonical trong của bạn own site — đây là không một claim đó noindex và canonical có thể không bao giờ technically coexist on một trang (một trang bạn là genuinely retiring có thể vẫn carry một self-referencing canonical). MỘT canonical pointing tại một khác nhau URL deserves đó strongest warning: xác nhận đó exclusion và consolidation là cả hai dự kiến. Dùng canonical để consolidate duplicates; dùng noindex chỉ khi bạn genuinely muốn này trang out of kết quả.

noindex,follow so với noindex,nofollow — chậm decay

MỘT phổ biến pattern là noindex,follow: giữ đó trang out of kết quả, nhưng giữ sau của nó links so equity vẫn luồng qua điều này (handy during một migration hoặc trong khi một trang là temporarily out). Hiện tại chính thức Google tài liệu không mô tả này decaying tự động — điều này explicitly cho phép combining noindex với other rules, including setting noindex,nofollow on purpose từ day một. Điều gì I’m relying on cho đó “it fades over time” (bản dịch) «điều này fades theo thời gian» claim là một 2017 quản trị viên web hangout, nơi John Mueller đã nói một dài-term noindex tends để end lên treated như noindex,nofollow trong thực tế: khi Google decides đó trang thực sự không belong trong tìm kiếm và drops điều này completely, điều này cũng dừng sau đó trang links, vì đây là đã dừng processing đó trang tại all. đó là một practitioner observation từ một video transcript, không một được ghi lại Google policy, so treat điều này as directional thay vì guaranteed. Either way, đó practical takeaway holds: noindex,follow là fine cho một transitional period, nhưng không lean on điều này as một vĩnh viễn link-equity strategy — plan để cách sửa đó underlying links (hoặc xóa đó trang) thay vì.

Cách dài làm noindex take?

Không instantly. noindex chỉ áp dụng sau Google recrawls và reprocesses đó trang — until thì, đó trang có thể stay được lập chỉ mục mặc dù đó tag là trực tiếp. Google không commit để một fixed window, và của nó own hướng dẫn leans toward “could be a while,” (bản dịch) «có thể là một trong khi,» không “bất kỳ day hiện tại”: “Depending on the importance of the page on the internet, it may take months for Googlebot to revisit a page.” (bản dịch) «Depending on đó importance of đó trang on đó internet, điều này có thể take months cho Googlebot để revisit một trang.» MỘT cao-traffic, frequently-linked trang có thể nhận recrawled trong days; một thấp-giá trị, rarely-linked một có thể sit cho months. Nếu bạn cần một trang out of kết quả urgently, đó GSC Removals tool là một stopgap (điều này hides đó URL temporarily trong khi đó vĩnh viễn noindex làm của nó chậm hơn hoạt động). Cho genuinely đã biến mất các trang, một 404/410 cũng drops them: as I wrote trong my removal hướng dẫn, “If you remove the page and serve either a 404 (not found) or 410 (gone) status code, then the page will be removed from the index shortly after the page is re-crawled.” (bản dịch) «Nếu bạn xóa đó trang và serve either một 404 (không tìm thấy) hoặc 410 (đã biến mất) mã trạng thái, thì đó trang sẽ bị gỡ bỏ từ đó chỉ mục shortly sau đó trang là re-được crawl.» Giống nhau theme mọi nơi — điều này happens on recrawl.

noindex trong robots.txt là dead (since Sept 1, 2019)

Bạn’ll vẫn see mọi người suggest một Noindex: line trong robots.txt. không. Điều này đã là không bao giờ an officially supported rule, và Google retired ngay cả của nó không chính thức xử lý năm ago. Từ đó July 2019 Tìm kiếm Central announcement: “Since these rules were never documented by Google, naturally, their usage in relation to Googlebot is very low.” (bản dịch) «Since những rules đã là không bao giờ được ghi lại by Google, naturally, của họ usage trong relation để Googlebot là very thấp.» Và đó date: “we’re retiring all code that handles unsupported and unpublished rules (such as noindex) on September 1, 2019.” (bản dịch) «chúng ta là retiring all code đó xử lý unsupported và unpublished rules (such as ) on September 1, 2019.»

Đó giống nhau post named đó supported alternatives, và noindex qua đó meta tag / header topped đó list: noindex in robots meta tags: Supported both in the HTTP response headers and in HTML, the noindex rule is the most effective way to remove URLs from the index when crawling is allowed.” (bản dịch) «trong robots meta tags: Supported cả hai trong đó HTTP các header phản hồi và trong HTML, đó rule là đó hầu hết effective way để xóa URLs từ đó chỉ mục khi crawling là được phép.» (Cũng listed: 404/410 status codes, password protection, robots.txt disallow cho crawl prevention, và đó Search Console removal tool.)

Cách verify noindex trong Google Search Console

Hai kiểm tra:

  • URL Inspection. Chạy đó URL qua Inspect, thì Kiểm thử trực tiếp URL. Điều này tells bạn liệu đó trang là indexable và liệu Google sees một noindex directive — đó fastest way để xác nhận đó tag là đang đọc on đó trực tiếp trang.
  • Trang Lập chỉ mục báo cáo. Noindexed các trang là listed dưới đó status “URL marked ‘noindex’” (bản dịch) «URL được đánh dấu noindex» trong đó Không được lập chỉ mục section. Google help text: “When Google tried to index the page it encountered a ‘noindex’ directive and therefore did not index it.” (bản dịch) «Khi Google tried để chỉ mục đó trang điều này encountered một ‘noindex’ directive và do đó đã không chỉ mục điều này.» Nếu đó là một trang bạn wanted được lập chỉ mục, đó là của bạn bug — xóa đó directive.

Một naming note cho anyone searching old ghi-ups: đó legacy Coverage báo cáo called này “Excluded by ‘noindex’ tag.” (bản dịch) «Bị loại trừ bởi thẻ noindex.» Đó hiện tại Trang Lập chỉ mục báo cáo dùng “URL marked ‘noindex’” (bản dịch) «URL được đánh dấu noindex» — giống nhau điều, newer label.

Điều gì noindex không bảo đảm

một vài điều mọi người assume noindex buys them đó nó thực ra không:

  • Crawl-budget savings. Google vẫn có để fetch trang để see tag — noindex alone không reduce crawling. nếu bạn muốn đó cũng, thêm disallow trong robots.txt, nhưng chỉ sau khi trang có đã dropped từ chỉ mục (see mistake trên cho Vì sao đang làm nó lên front backfires).
  • Instant removal. Covered trên — nó happens on recrawl, với không fixed timetable, và Google itself nói thấp hơn-priority trang có thể take months.
  • Duplicate consolidation. đó Điều gì rel="canonical" là cho; noindex chỉ xóa trang từ Tìm kiếm, nó không hợp nhất các tín hiệu toward một URL.
  • Confidentiality. trang vẫn giữ publicly requestable by anyone với URL. nếu điều gì đó thực ra cần để là riêng tư, đó authentication vấn đề, không tìm kiếm-directive vấn đề.
  • Xếp hạng recovery nếu bạn reverse nó. Removing noindex không restore trang old thứ hạng — Google có để recrawl, re-evaluate, và effectively re-earn của nó position từ scratch.
  • Giống hệt timing trên các công cụ tìm kiếm. Bing và khác engines chạy của họ own crawl và recrawl schedules independently của Google.
  • Exclusion từ mỗi non-tìm kiếm sử dụng của bạn nội dung. noindex chặn trang từ Google Search as toàn bộ — including Tìm kiếm own AI features (AI Overviews và similar draw trên các trang đó là được lập chỉ mục và eligible để là shown, so noindexed trang là out của những điều đó cũng). Điều gì nó làm không làm là control Google tách biệt Google-Extended setting, mà governs liệu của bạn nội dung có thể là được sử dụng để train hoặc ground Google generative AI models bên ngoài của Tìm kiếm. những điều đó là hai khác controls cho hai khác jobs.

nơi noindex fits với mọi thứ khác

noindex là lever bạn reach cho Khi trang là trong chỉ mục nhưng không nên là — cure cho một flavor của chỉ mục bloat (thin, utility, hoặc duplicate-ish các trang với không tìm kiếm giá trị). nó sits right tiếp theo để robots meta tag và X-Robots-Tag header (của nó hai phân phối các phương thức), robots.txt và của nó disallow directive ( crawl control nó so thường confused với), thẻ canonical (sử dụng đó cho duplicate consolidation, không noindex), và rộng hơn crawling và lập chỉ mục stages nó plugs vào. Nhận crawl-so với-chỉ mục phân biệt right và noindex dừng là mysterious: cho phép crawl, phục vụ tag, chờ cho recrawl.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.