indeksowanie

How wyszukiwarki sklep i organize strony so they może rank — treść analiza, canonicalization, why crawled isn't zindeksowany, i reading the GSC strona indeksowanie raport.

Opublikowano po raz pierwszy: 23 cze 2026 · Ostatnia aktualizacja: 3 sie 2026 · Advanced
Języki
1 sygnał dowodowy na tej stronie

indeksowanie jest stage two of search (crawl → index → serve): po a strona jest crawled, the engine understands it, deduplicates i canonicalizes it, i — if it qualifies — sklepy it in the search index. Crawled isn't zindeksowany; Google selects co to zachować, i indeksowanie isn't guaranteed. It's nie a czynnik rankingowy, ale a strona musi być zindeksowany przed it może rank. To zachować a strona out, używać noindex i leave it crawlable — don't blok it in robots.txt. ten hub wyjaśnia the whole stage i routes you to the deep dives.

TL;DR — indeksowanie jest the second of search’s three stages (crawl → index → serve): Google understands a crawled strona (tekst, key znaczniki, images, video; it renders JS), detects duplicates, clusters similar strony i picks the najbardziej representative one (canonicalization — rel=canonical jest a hint, nie a reguła), computes signals, i sklepy the canonical in the index. Crawled ≠ zindeksowany — “indexing isn’t guaranteed,” i the call jest largely o quality/wartość. The Search Console strona indeksowanie raport jest twój cockpit. To deindex, używać noindex i zachowaj strona crawlable; nigdy używać robots.txt to remove a strona, ponieważ a blocked strona może nadal być zindeksowany (just bez a snippet).

Evidence for this claim Google must be able to crawl a page to see and apply its noindex rule. Scope: Google-supported robots meta and X-Robots-Tag directives; a robots.txt block can prevent Google from seeing the rule. Confidence: high · Verified: Google Search Central: Block Search indexing with noindex

indeksowanie jest stage two of three

Indexing is stage two of three — the middle filter between crawling and ranking. Źródło: /technical-seo/how-search-works/indexing/

Three stages run left to right. Crawl discovers and downloads a URL. Index processes the page and stores eligible information. Serve or rank orders the best indexed matches for a query. The Index stage is highlighted, and a note says not every page advances through every stage.

© Patrick Stox LLC · CC BY 4.0 ·

Google jest blunt o the pipeline: “Google Search works in three stages, and not all pages make it through each stage” — crawling, indeksowanie, i serving. indeksowanie jest the middle stage, i the doc defines it cleanly: “Indexing: Google analyzes the text, images, and video files on the page, and stores the information in the Google index, which is a large database.”

Evidence for this claim Google Search describes crawling, indexing, and serving as three distinct stages; indexing analyzes page content and stores eligible information in Google's index. Scope: web search Confidence: high · Verified: In-depth guide to how Google Search works

A strona ma to być crawled przed it może być zindeksowany, i it ma to być zindeksowany przed it może rank. ale none of tamte są guarantees — każdy stage jest a filter. Keeping the three stages oddzielny in twój head jest the single najbardziej użyteczny mental model in techniczne SEO, i it’s why I zawsze ask który stage a strona jest failing at przed changing anything. (dla the stage przed ten one, see the crawling hub — crawl → index jest the pipeline.)

co actually happens podczas indeksowanie

Indexing is a sequence: understand, cluster and select a canonical, then store. Źródło: /technical-seo/how-search-works/indexing/

Step one analyzes a crawled page for text, title, alt text, images, and video. Step two groups duplicate URLs into a cluster and chooses the most representative page as canonical. Step three stores the canonical page and its cluster information in the Google index.

© Patrick Stox LLC · CC BY 4.0 ·

indeksowanie isn’t one thing; it’s a sequence:

  • Understanding the treść. Google: “After a page is crawled, Google tries to understand what the page is about. This stage is called indexing.” że means “Google analyzes the textual content and key content tags and attributes, such as <title> elements and alt attributes, images, videos, and more.” JavaScript jest wyrenderowany as part of ten — if twój treść tylko appears po JS runs, it nadal ma to render przed it może być understood.
  • Duplicate detection & canonicalization. ten jest the part najbardziej explainers skip, i it’s gdzie a lot of “why isn’t this indexed?” mysteries live. Google “determines if a page is a duplicate of another page on the internet or canonical.” The mechanic: “we first group together (also known as clustering) the pages that we found on the internet that have similar content, and then we select the one that’s most representative of the group.” że representative jest the canonical — “The canonical is the page that may be shown in search results.”
  • Computing signals & storing. Finally, “The collected information about the canonical page and its cluster may be stored in the Google index, a large database hosted on thousands of computers.” Google’s named indeksowanie system behind wszystkie ten jest Caffeine — the warstwa że ingests crawl data, renders i extracts, computes signals, i builds the index że gets served.

Canonicalization: a hint, nie a command

ponieważ canonicalization happens podczas indeksowanie, it deserves jego own note. “Canonicalization is the process of selecting the representative –canonical– URL of a piece of content,” i it exists ponieważ “this process helps Google show only one version of the otherwise duplicate content in its search results.”

The load-bearing detail: twój rel=canonical jest a suggestion. Google’s words: “indicating a canonical preference is a hint, not a rule.” Google weighs wiele signals — in my canonicalization poradnik I note że, per Google’s Allan Scott, there są roughly 40 różny canonical selection signals — i it może pick a różny URL than the one you flagged. że’s exactly co the “Duplicate, Google chose different canonical than user” status in Search Console jest telling you.

Crawled ≠ zindeksowany: why strony don’t get zindeksowany

Here’s the myth-buster, straight z the docs: “Indexing isn’t guaranteed; not every page that Google processes will be indexed.” Evidence for this claim Google does not guarantee that every processed page will be indexed. Scope: Google Search indexing; the source gives examples of possible causes rather than an exhaustive decision formula. Confidence: high · Verified: Google Search Central: In-depth guide to how Google Search works Google listy common powody it fails — “The quality of the content on page is low,” “Robots meta rules disallow indexing,” i “The design of the website might make indexing difficult.”

The reps są even więcej bezpośredni że ten jest a selection decision driven by wartość, nie a quota you może kupić past:

  • John Mueller, on how long “Discovered/Crawled – currently not indexed” może persist: “That can be forever. It’s something where we just don’t crawl and index all pages.” The fix isn’t resubmitting — it’s making the systemy recognize the wartość, to “continue working on the website and making sure that our systems recognize that there’s value in crawling and indexing more and then over time we will crawl and index more.”
  • Mueller again: “it’s important to keep in mind that Google just doesn’t index every page on the web, even if it’s submitted directly.” i, bluntly: “Well, lots of SEOs & sites (perhaps not you/yours!) produce terrible content that’s not worth indexing.”
  • Gary Illyes, on why it’s selective: “we don’t have infinite space, so we want to index stuff that we think– well, not we– but our algorithms determine that it might be searched for…”
  • Martin Splitt frames it as a balancing act: “I usually describe it as a challenge with the balance between not overwhelming the website and also spending our resources where it matters.”

The practical takeaway: a mapa witryny lub “request indexing” aids discovery, nie selection. Submitting a strona again won’t force it in. The lever jest witryna quality i wartość.

Reading the Google Search Console strona indeksowanie raport

The strona indeksowanie raport jest gdzie indeksowanie problems actually pokazywać up. Treat każdy status as a diagnosis. te są Google’s own verbatim opisy:

  • Crawled – currently nie zindeksowany: “The page was crawled by Google but not indexed. It may or may not be indexed in the future; no need to resubmit this URL for crawling.” zwykle a quality/wartość judgment — poprawić the strona, don’t spam the resubmit button.
  • odkryty – currently nie zindeksowany: “The page was found by Google, but not crawled yet. Typically, Google wanted to crawl the URL but this was expected to overload the site; therefore Google rescheduled the crawl.” Technically a pre-crawl, capacity-driven status — ale if it persists, reps tie że to wartość, same as the one above.
  • Duplicate bez użytkownik-wybrany canonical: “This page is a duplicate of another page, although it doesn’t indicate a preferred canonical page. Google has chosen the other page as the canonical for this page, and so will not serve this page in Search.”
  • Duplicate, Google chose różny canonical than użytkownik: “This page is marked as canonical for a set of pages, but Google thinks another URL makes a better canonical.” (The “hint, not a rule” outcome in the wild.)
  • Alternate strona z proper znacznik kanoniczny: “This page is marked as an alternate of another page… This page correctly points to the canonical page, which is indexed, so there is nothing you need to do.”
  • zindeksowany, though blocked by robots.txt: “The page was indexed despite being blocked by your website’s robots.txt file. Google always respects robots.txt, but this doesn’t necessarily prevent indexing if someone else links to your page.” ten jest the proof że blocking crawling robi nie blok indeksowanie.
  • URL blocked by robots.txt: “This page was blocked by your site’s robots.txt file.”
  • URL marked ‘noindex’: “When Google tried to index the page it encountered a ‘noindex’ directive and therefore did not index it.” (ten jest the deindex działający as intended.)
  • strona z redirect: “This is a non-canonical URL that redirects to another page. As such, this URL will not be indexed.”
  • Soft 404: “The page request returns what we think is a soft 404 response. This means that it returns a user-friendly ‘not found’ message but not a 404 HTTP response code.”

How to control indeksowanie the right way

To get a strona zindeksowany: make it crawlable, link to it internally, obejmować it in twój mapa witryny — i, above wszystkie, make it worth indeksowanie. Discovery aids don’t override the wartość judgment.

To zachować a strona OUT — the najbardziej-botched control in SEO: używać noindex, “a rule set with either a <meta> tag or HTTP response header,” i zachowaj strona crawlable. Google’s load-bearing ostrzeżenie: “For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler.”

Evidence for this claim For Google to apply noindex, the crawler must be allowed to access the page or resource; a robots.txt block can prevent Google from seeing the rule. Scope: HTML and HTTP resources Confidence: high · Verified: Block Search indexing with noindex

The mistake I see constantly jest adding noindex i blocking the strona in robots.txt. że’s counterproductive. As I put it in How to Remove URLs z Google Search: “For these tags to be seen, a search engine needs to be able to crawl the pages—so make sure they aren’t blocked in robots.txt,” i “Crawling is not the same thing as indexing. Even if Google is blocked from crawling pages, if there are any internal or external links to a page they can still index it.” In my piece on the zindeksowany, though blocked by robots.txt status I say it even więcej plainly: “Unless Google can crawl a page, they won’t see the noindex meta tag and may still index it because it has links.” The fix: “Just add a noindex meta robots tag and make sure to allow crawling—assuming it’s canonical.”

dla the fuller removal decision tree — 404/410 vs. noindex, the Removals narzędzie’s ~6-month hold, i password protection — see how to deindex a strona.

indeksowanie in Bing

Bing runs the same pipeline. As Microsoft describes it: “As Bingbot crawls the web, it sends information to Bing about what it finds. These pages are then added to the Bing index.” The same controls apply — a noindex directive zachowuje a strona out, i an ponad-restrictive robots.txt może stop Bingbot z ever crawling it. Bing również needs co najmniej one link pointing to twój witryna to find it in the pierwszy place.

nie każdy strona belongs in the index

A prosty filter, nie a universal reguła: a strona jest worth indeksowanie if it może pokazywać up dla a search z a distinct, użyteczny wynik. że’s the bar to sprawdzenie duplicates, parametr warianty, private lub staging URLs, i thin lub repetitive stan magazynowy strony wobec — nie a powód to noindex a whole strona type by domyślny. Run the pełny audit on the strony built to own it, następny.

gdzie to go następny: the indeksowanie cluster

ten hub jest the overview. Two things go błędny at scale, i każdy gets jego own deep dive:

  • Index bloat — gdy too wiele niski-wartość, duplicate, lub thin URLs end up in the index, diluting twój witryna i wasting crawl/index zasoby. How to diagnose it i prune it safely.
  • mobilny-pierwszy indeksowanie — Google indexes the mobilny version of twój strony, so treść, links, i dane strukturalne mieć to reach parity między mobilny i komputer stacjonarny. co to sprawdzenie i co breaks.

oba topics są nested poniżej ten hub — they’re in the sidebar too, i they’ll link back here.

The stage przed ten one — how bots odkrywać i download twój strony — lives in the crawling hub; crawl → index jest the pipeline, i a strona ma to jasny crawling przed dowolny of ten applies. dla the broader picture, see How Search działa.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.