Lập chỉ mục

Cách các công cụ tìm kiếm store và organize các trang so they có thể xếp hạng — nội dung analysis, canonicalization, vì sao được crawl không được lập chỉ mục, và reading đó GSC Trang lập chỉ mục báo cáo.

Xuất bản lần đầu: 23 thg 6, 2026 · Cập nhật lần cuối: 8 thg 8, 2026 · Advanced
Ngôn ngữ
1 tín hiệu bằng chứng trên trang này

Lập chỉ mục là giai đoạn hai of tìm kiếm (crawl → chỉ mục → serve): sau một trang là được crawl, đó engine understands điều này, deduplicates và canonicalizes điều này, và — nếu điều này qualifies — stores điều này trong đó tìm kiếm chỉ mục. Được crawl không được lập chỉ mục; Google selects điều cần giữ, và lập chỉ mục không guaranteed. đây là không phải là yếu tố xếp hạng, nhưng một trang phải được lập chỉ mục trước điều này có thể xếp hạng. Để giữ một trang out, dùng noindex và leave điều này crawlable — không block điều này trong robots.txt. Này hub giải thích đó toàn bộ stage và routes bạn để đó deep dives.

TL;DR — Lập chỉ mục là đó second of tìm kiếm three stages (crawl → chỉ mục → serve): Google understands một được crawl trang (text, key tags, images, video; điều này renders JS), detects duplicates, clusters similar các trang và picks đó hầu hết representative một (canonicalization — rel=canonical là một hint, không một rule), computes các tín hiệu, và stores đó canonical trong đó chỉ mục. Được crawl ≠ được lập chỉ mục — “indexing isn’t guaranteed,” (bản dịch) «lập chỉ mục không guaranteed,» và đó call là largely về quality/giá trị. Đó Tìm kiếm Console Trang lập chỉ mục báo cáo là của bạn cockpit. Để deindex, dùng noindex và giữ đó trang crawlable; không bao giờ dùng robots.txt để xóa một trang, vì một blocked trang có thể vẫn là được lập chỉ mục (chỉ không có một snippet).

Evidence for this claim Google must be able to crawl a page to see and apply its noindex rule. Scope: Google-supported robots meta and X-Robots-Tag directives; a robots.txt block can prevent Google from seeing the rule. Confidence: high · Verified: Google Search Central: Block Search indexing with noindex

lập chỉ mục là giai đoạn hai của three

Indexing is stage two of three — the middle filter between crawling and ranking. Nguồn: /technical-seo/how-search-works/indexing/

Three stages run left to right. Crawl discovers and downloads a URL. Index processes the page and stores eligible information. Serve or rank orders the best indexed matches for a query. The Index stage is highlighted, and a note says not every page advances through every stage.

© Patrick Stox LLC · CC BY 4.0 ·

Google là blunt về đó pipeline: “Google Search works in three stages, and not all pages make it through each stage” (bản dịch) «Google Search hoạt động trong three stages, và không all các trang làm điều này qua mỗi stage» — crawling, lập chỉ mục, và serving. Lập chỉ mục là đó middle stage, và đó doc defines điều này cleanly: “Indexing: Google analyzes the text, images, and video files on the page, and stores the information in the Google index, which is a large database.” (bản dịch) «Lập chỉ mục: Google analyzes đó text, images, và video files on đó trang, và stores đó information trong đó Google chỉ mục, mà là một lớn database.»

Evidence for this claim Google Search describes crawling, indexing, and serving as three distinct stages; indexing analyzes page content and stores eligible information in Google's index. Scope: web search Confidence: high · Verified: In-depth guide to how Google Search works

trang có để là được crawl trước khi nó có thể là được lập chỉ mục, và nó có để là được lập chỉ mục trước khi nó có thể xếp hạng. nhưng none của những điều đó là bảo đảm — mỗi stage là filter. Giữ three stages tách biệt trong của bạn head là single phần lớn hữu ích mental model trong kỹ thuật SEO, và nó Vì sao I luôn ask stage trang là failing tại trước khi thay đổi bất cứ điều gì. (cho stage trước khi điều này một, see crawling hub — crawl → chỉ mục là pipeline.)

Điều gì thực ra happens during lập chỉ mục

Indexing is a sequence: understand, cluster and select a canonical, then store. Nguồn: /technical-seo/how-search-works/indexing/

Step one analyzes a crawled page for text, title, alt text, images, and video. Step two groups duplicate URLs into a cluster and chooses the most representative page as canonical. Step three stores the canonical page and its cluster information in the Google index.

© Patrick Stox LLC · CC BY 4.0 ·

lập chỉ mục không phải một điều; nó sequence:

  • Understanding đó nội dung. Google: “After a page is crawled, Google tries to understand what the page is about. This stage is called indexing.” (bản dịch) «Sau một trang là được crawl, Google tries để understand điều gì đó trang là về. Này stage là called lập chỉ mục.» Đó có nghĩa là “Google analyzes the textual content and key content tags and attributes, such as <title> elements and alt attributes, images, videos, and more.” (bản dịch) «Google analyzes đó textual nội dung và key nội dung tags và các thuộc tính, such as <title> elements và alt các thuộc tính, images, videos, và hơn.» JavaScript là được kết xuất as part of này — nếu nội dung của bạn chỉ xuất hiện sau JS chạy, điều này vẫn có để render trước điều này có thể là understood.
  • Duplicate detection & canonicalization. Này là đó part hầu hết explainers skip, và đây là nơi một lot of “why isn’t this indexed?” (bản dịch) «vì sao không này được lập chỉ mục?» mysteries trực tiếp. Google “determines if a page is a duplicate of another page on the internet or canonical.” (bản dịch) «determines nếu một trang là một duplicate of một sản phẩm khác trang on đó internet hoặc canonical.» Đó mechanic: “we first group together (also known as clustering) the pages that we found on the internet that have similar content, and then we select the one that’s most representative of the group.” (bản dịch) «we đầu tiên group together (cũng known as clustering) đó các trang đó we được tìm thấy on đó internet đó có similar nội dung, và thì we select đó một đó là hầu hết representative of đó group.» Đó representative là đó canonical — “The canonical is the page that may be shown in search results.” (bản dịch) «Đó canonical là đó trang đó có thể là shown trong kết quả tìm kiếm.»
  • Computing các tín hiệu & storing. Finally, “The collected information about the canonical page and its cluster may be stored in the Google index, a large database hosted on thousands of computers.” (bản dịch) «Đó collected information về đó canonical trang và của nó cluster có thể là stored trong đó Google chỉ mục, một lớn database hosted on thousands of computers.» Google named hệ thống lập chỉ mục behind all này là Caffeine — đó layer đó ingests crawl dữ liệu, renders và extracts, computes các tín hiệu, và xây dựng đó chỉ mục đó nhận phân phối.

Canonicalization: hint, không command

Vì canonicalization happens during lập chỉ mục, điều này deserves của nó own note. “Canonicalization is the process of selecting the representative –canonical– URL of a piece of content,” (bản dịch) «Canonicalization là đó xử lý of selecting đó representative –canonical– URL of một piece of nội dung,» và điều này tồn tại vì “this process helps Google show only one version of the otherwise duplicate content in its search results.” (bản dịch) «này xử lý helps Google cho thấy chỉ một version of đó nếu không duplicate nội dung trong của nó kết quả tìm kiếm.»

Đó load-bearing detail: của bạn rel=canonical là một suggestion. Google words: “indicating a canonical preference is a hint, not a rule.” (bản dịch) «indicating một canonical preference là một hint, không một rule.» Google weighs nhiều các tín hiệu — trong my canonicalization hướng dẫn I note đó, theo Google Allan Scott, có khoảng 40 khác nhau canonical selection các tín hiệu — và điều này có thể pick một khác nhau URL hơn đó một bạn flagged. đó là chính xác điều gì đó “Duplicate, Google chose different canonical than user” (bản dịch) «Duplicate, Google chose khác nhau canonical hơn người dùng» status trong Search Console là telling bạn.

Được crawl ≠ được lập chỉ mục: Vì sao các trang không nhận được lập chỉ mục

Ở đây đó myth-buster, straight từ đó tài liệu: “Indexing isn’t guaranteed; not every page that Google processes will be indexed.” (bản dịch) «Lập chỉ mục không guaranteed; không mỗi trang đó Google xử lý sẽ là được lập chỉ mục.» Evidence for this claim Google does not guarantee that every processed page will be indexed. Scope: Google Search indexing; the source gives examples of possible causes rather than an exhaustive decision formula. Confidence: high · Verified: Google Search Central: In-depth guide to how Google Search works Google lists phổ biến reasons điều này fails — “The quality of the content on page is low,” (bản dịch) «Đó quality of đó nội dung on trang là thấp,» “Robots meta rules disallow indexing,” (bản dịch) «Robots meta rules disallow lập chỉ mục,»“The design of the website might make indexing difficult.” (bản dịch) «Đó design of đó website có thể làm lập chỉ mục difficult.»

reps là ngay cả nhiều hơn trực tiếp đó Đây là selection decision driven by giá trị, không quota Bạn có thể buy past:

  • John Mueller, on cách dài “Discovered/Crawled – currently not indexed” (bản dịch) «Discovered/Được crawl – hiện tại không được lập chỉ mục» có thể persist: “That can be forever. It’s something where we just don’t crawl and index all pages.” (bản dịch) «Điều đó có thể kéo dài mãi. Có những trường hợp chúng tôi chỉ không thu thập dữ liệu và lập chỉ mục mọi trang.» Đó cách sửa không resubmitting — đây là đang làm đó các hệ thống recognize đó giá trị, để “continue working on the website and making sure that our systems recognize that there’s value in crawling and indexing more and then over time we will crawl and index more.” (bản dịch) «tiếp tục cải thiện trang web và bảo đảm hệ thống nhận ra giá trị của việc thu thập dữ liệu và lập chỉ mục nhiều hơn; theo thời gian, chúng tôi sẽ thu thập dữ liệu và lập chỉ mục nhiều hơn.»
  • Mueller again: “it’s important to keep in mind that Google just doesn’t index every page on the web, even if it’s submitted directly.” (bản dịch) «đây là quan trọng để hãy nhớ rằng đó Google chỉ không chỉ mục mỗi trang on đó web, ngay cả khi đây là được gửi trực tiếp.» Và, bluntly: “Well, lots of SEOs & sites (perhaps not you/yours!) produce terrible content that’s not worth indexing.” (bản dịch) «Well, lots of SEOs & các trang (perhaps không bạn/yours!) produce terrible nội dung đó là không worth lập chỉ mục.»
  • Gary Illyes, on vì sao đây là selective: “we don’t have infinite space, so we want to index stuff that we think– well, not we– but our algorithms determine that it might be searched for…” (bản dịch) «we không có infinite space, so we muốn để chỉ mục stuff đó we think– well, không we– nhưng của chúng ta algorithms determine đó điều này có thể là searched cho…»
  • Martin Splitt frames điều này as một balancing act: “I usually describe it as a challenge with the balance between not overwhelming the website and also spending our resources where it matters.” (bản dịch) «I thường mô tả điều này as một challenge với đó balance giữa không overwhelming đó website và cũng spending của chúng ta các tài nguyên nơi điều này matters.»

Đó practical takeaway: một sitemap hoặc “request indexing” (bản dịch) «yêu cầu lập chỉ mục» aids phát hiện, không selection. Submitting một trang again sẽ không force điều này trong. Đó lever là site quality và giá trị.

Reading Google Search Console trang lập chỉ mục báo cáo

trang lập chỉ mục báo cáo là nơi lập chỉ mục các vấn đề thực ra hiển thị lên. Treat mỗi status as diagnosis. những điều này là Google own verbatim các mô tả:

  • Được crawl – hiện tại không được lập chỉ mục: “The page was crawled by Google but not indexed. It may or may not be indexed in the future; no need to resubmit this URL for crawling.” (bản dịch) «Đó trang đã là được crawl by Google nhưng không được lập chỉ mục. Điều này có thể hoặc có thể không là được lập chỉ mục trong đó tương lai; không cần để resubmit này URL cho crawling.» Thường một quality/giá trị judgment — improve đó trang, không spam đó resubmit button.
  • Discovered – hiện tại không được lập chỉ mục: “The page was found by Google, but not crawled yet. Typically, Google wanted to crawl the URL but this was expected to overload the site; therefore Google rescheduled the crawl.” (bản dịch) «Đó trang đã là được tìm thấy by Google, nhưng không được crawl tuy vậy. Thông thường, Google muốn crawl đó URL nhưng điều này được dự kiến sẽ làm quá tải trang web; làm đó Google đã lên lịch lại lượt crawl.» Technically một pre-crawl, capacity-driven status — nhưng nếu điều này persists, reps tie đó để giá trị, giống nhau as đó một trên.
  • Duplicate không có người dùng-được chọn canonical: “This page is a duplicate of another page, although it doesn’t indicate a preferred canonical page. Google has chosen the other page as the canonical for this page, and so will not serve this page in Search.” (bản dịch) «Này trang là một duplicate of một sản phẩm khác trang, although điều này không indicate một được ưu tiên canonical trang. Google có chosen đó other trang as đó canonical cho này trang, và so sẽ không serve này trang trong Tìm kiếm.»
  • Duplicate, Google chose khác nhau canonical hơn người dùng: “This page is marked as canonical for a set of pages, but Google thinks another URL makes a better canonical.” (bản dịch) «Này trang là marked as canonical cho một set of các trang, nhưng Google thinks một sản phẩm khác URL làm một tốt hơn canonical.» (Đó “hint, not a rule” (bản dịch) «hint, không một rule» outcome trong đó wild.)
  • Alternate trang với proper canonical tag: “This page is marked as an alternate of another page… This page correctly points to the canonical page, which is indexed, so there is nothing you need to do.” (bản dịch) «Này trang là marked as an alternate of một sản phẩm khác trang… Này trang correctly points để đó canonical trang, mà là được lập chỉ mục, so có không có gì bạn cần để làm.»
  • Được lập chỉ mục, though blocked by robots.txt: “The page was indexed despite being blocked by your website’s robots.txt file. Google always respects robots.txt, but this doesn’t necessarily prevent indexing if someone else links to your page.” (bản dịch) «Đó trang đã là được lập chỉ mục despite đang blocked by của bạn website robots.txt file. Google luôn respects robots.txt, nhưng này không nhất thiết ngăn lập chỉ mục nếu ai đó khác links để trang của bạn.» Này là đó proof đó blocking crawling không block lập chỉ mục.
  • URL blocked by robots.txt: “This page was blocked by your site’s robots.txt file.” (bản dịch) «Này trang đã là blocked by trang web của bạn robots.txt file.»
  • URL marked ‘noindex’: “When Google tried to index the page it encountered a ‘noindex’ directive and therefore did not index it.” (bản dịch) «Khi Google tried để chỉ mục đó trang điều này encountered một ‘noindex’ directive và do đó đã không chỉ mục điều này.» (Này là đó deindex hoạt động as dự kiến.)
  • Trang với chuyển hướng: “This is a non-canonical URL that redirects to another page. As such, this URL will not be indexed.” (bản dịch) «Này là một non-canonical URL đó các chuyển hướng để một sản phẩm khác trang. As such, này URL sẽ không là được lập chỉ mục.»
  • Soft 404: “The page request returns what we think is a soft 404 response. This means that it returns a user-friendly ‘not found’ message but not a 404 HTTP response code.” (bản dịch) «Đó trang yêu cầu trả về điều gì we think là một soft 404 phản hồi. Này có nghĩa là đó điều này trả về một người dùng-friendly ‘không tìm thấy’ message nhưng không một 404 HTTP phản hồi code.»

Cách control lập chỉ mục right way

để nhận trang được lập chỉ mục: làm nó crawlable, link để nó internally, bao gồm nó trong của bạn sitemap — và, trên all, làm nó worth lập chỉ mục. Phát hiện aids không override giá trị judgment.

Để giữ một trang OUT — đó hầu hết-botched control trong SEO: dùng noindex, “a rule set with either a <meta> tag or HTTP response header,” (bản dịch) «một rule set với either một <meta> tag hoặc HTTP header phản hồi,»giữ đó trang crawlable. Google load-bearing warning: “For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler.” (bản dịch) «Cho đó noindex rule để là effective, đó trang hoặc tài nguyên không được là blocked by một robots.txt file, và điều này có để là nếu không accessible để đó crawler.»

Evidence for this claim For Google to apply noindex, the crawler must be allowed to access the page or resource; a robots.txt block can prevent Google from seeing the rule. Scope: HTML and HTTP resources Confidence: high · Verified: Block Search indexing with noindex

Đó mistake I see constantly là thêm noindex blocking đó trang trong robots.txt. đó là counterproductive. As I put điều này trong Cách Xóa URLs Từ Google Search: “For these tags to be seen, a search engine needs to be able to crawl the pages—so make sure they aren’t blocked in robots.txt,” (bản dịch) «Cho những tags để là seen, một công cụ tìm kiếm cần để là able để crawl đó các trang—so hãy bảo đảm they không blocked trong robots.txt,» và “Crawling is not the same thing as indexing. Even if Google is blocked from crawling pages, if there are any internal or external links to a page they can still index it.” (bản dịch) «Crawling không phải cùng một điều as lập chỉ mục. Ngay cả khi Google là blocked từ crawling các trang, nếu có bất kỳ internal hoặc liên kết bên ngoài để một trang they có thể vẫn chỉ mục điều này.» Trong my piece on đó Được lập chỉ mục, though blocked by robots.txt status I chẳng hạn điều này ngay cả hơn plainly: “Unless Google can crawl a page, they won’t see the noindex meta tag and may still index it because it has links.” (bản dịch) «Trừ khi Google có thể crawl một trang, they sẽ không see đó noindex meta tag và có thể vẫn chỉ mục điều này vì điều này có links.» Đó cách sửa: “Just add a noindex meta robots tag and make sure to allow crawling—assuming it’s canonical.” (bản dịch) «Chỉ thêm một noindex meta robots tag và hãy bảo đảm để cho phép crawling—assuming đây là canonical.»

cho fuller removal cây quyết định — 404/410 so với. noindex, Removals tool ~6-month hold, và password protection — see Cách deindex trang.

lập chỉ mục trong Bing

Bing chạy đó giống nhau pipeline. As Microsoft mô tả điều này: “As Bingbot crawls the web, it sends information to Bing about what it finds. These pages are then added to the Bing index.” (bản dịch) «As Bingbot crawl đó web, điều này gửi information để Bing về điều gì điều này tìm thấy. Những các trang là thì đã thêm để đó Bing chỉ mục.» Đó giống nhau controls apply — một noindex directive giữ một trang out, và an over-restrictive robots.txt có thể dừng Bingbot từ bao giờ crawling điều này. Bing cũng cần ít nhất một link pointing để trang web của bạn để tìm điều này ngay từ đầu.

không mỗi trang belongs trong chỉ mục

đơn giản filter, không universal rule: trang là worth lập chỉ mục nếu nó có thể hiển thị lên cho tìm kiếm với distinct, hữu ích kết quả. đó bar để kiểm tra duplicates, parameter variants, riêng tư hoặc staging các URL, và thin hoặc repetitive inventory các trang so với — không reason để noindex toàn bộ trang loại theo mặc định. Chạy đầy đủ audit on các trang được xây dựng để own nó, tiếp theo.

nơi để go tiếp theo: lập chỉ mục cluster

điều này hub là overview. Hai điều go sai tại quy mô, và mỗi nhận của nó own deep dive:

  • chỉ mục bloat — Khi cũng nhiều thấp-giá trị, duplicate, hoặc thin các URL end lên trong chỉ mục, diluting của bạn trang web và wasting crawl/chỉ mục các tài nguyên. Cách chẩn đoán nó và prune nó safely.
  • Mobile-đầu tiên lập chỉ mục — Google indexes mobile version của bạn các trang, so nội dung, links, và dữ liệu có cấu trúc có để reach parity giữa mobile và desktop. Điều gì để kiểm tra và Điều gì breaks.

Cả hai topics là nested dưới điều này hub — họ’re trong sidebar cũng, và họ’ll link lại ở đây.

stage trước khi điều này một — Cách bots discover và download của bạn các trang — lives trong crawling hub; crawl → chỉ mục là pipeline, và trang có để clear crawling trước khi bất kỳ của điều này áp dụng. cho rộng hơn picture, see Cách Tìm kiếm Hoạt động.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.