Robots.txt

Co naprawdę robi robots.txt — kontroluje pobieranie, a nie indeksowanie — oraz dokładna składnia, sposób obsługi przez Google i błędy, które mogą zaszkodzić witrynie.

Opublikowano po raz pierwszy: 23 cze 2026 · Ostatnia aktualizacja: 11 sie 2026 · Advanced
Języki
1 sygnał dowodowy na tej stronie

Robots.txt to zwykły plik tekstowy w katalogu głównym każdego hosta, który wskazuje robotom, jakie adresy URL mogą pobierać. Kontroluje pobieranie, a nie indeksowanie: zablokowany adres nadal może znaleźć się w indeksie bez opisu. Do usunięcia strony z indeksu służy noindex, a strona nie może być wtedy zablokowana w robots.txt. Google obsługuje user-agent, allow, disallow i sitemap; plik ma limit 500 KiB, jest zwykle buforowany około 24 godzin, odpowiedź 4xx zwykle oznacza brak ograniczeń, a błędy 5xx mogą zatrzymać pobieranie całej witryny. Plik jest publiczny i nie chroni poufnych treści.

TL;DR — Robots.txt to zwykły plik tekstowy w katalogu głównym każdego hosta (/robots.txt, małymi literami), zgodny z protokołem Robots Exclusion Protocol (RFC 9309). Kontroluje pobieranie, nie indeksowanie. Google obsługuje tylko user-agent, allow, disallow i sitemap. Zakres obejmuje jeden host, protokół i port. Wygrywa najdłuższa, najbardziej szczegółowa reguła, a przy remisie najmniej restrykcyjna; * i $ są symbolami wieloznacznymi, a wielkość liter w ścieżkach ma znaczenie. Limit Google wynosi 500 KiB, pamięć podręczna zwykle działa około 24 godzin, 4xx poza 429 oznacza brak ograniczeń, a 5xx wstrzymuje pobieranie na około 12 godzin i uruchamia ostatnią poprawną kopię na około 30 dni. Nie blokuj CSS/JS potrzebnych do renderowania i nie traktuj publicznego pliku jako kontroli dostępu.

Czym jest robots.txt i gdzie się znajduje

Robots.txt realizuje Robots Exclusion Protocol, utworzony przez Martijna Kostera w 1994 r. i ustandaryzowany w 2022 r. jako RFC 9309, którego współautorami są Gary Illyes, Henner Zeller, Lizzi Sassman i Koster. “This document specifies and extends the ‘Robots Exclusion Protocol’ method originally defined by Martijn Koster in 1994 for service owners to control how content served by their services may be accessed, if at all, by automatic clients known as crawlers.” (tłumaczenie) „Dokument określa i rozszerza metodę Robots Exclusion Protocol zdefiniowaną pierwotnie przez Martijna Kostera w 1994 r., aby właściciele usług mogli kontrolować dostęp robotów do udostępnianych treści”.

Kilka ważnych faktów:

  • Plik musi znajdować się w katalogu głównym i mieć nazwę zapisaną małymi literami. “The rules MUST be accessible in a file named ‘/robots.txt’ (all lowercase) in the top-level path of the service.” (tłumaczenie) „Reguły MUSZĄ być dostępne w pliku /robots.txt zapisanym małymi literami w głównej ścieżce usługi”. Sam adres URL rozróżnia wielkość liter.
  • Zakres to jeden host, protokół i port. “The rules listed in the robots.txt file apply only to the host, protocol, and port number where the robots.txt file is hosted.” (tłumaczenie) „Reguły dotyczą wyłącznie hosta, protokołu i portu, na których udostępniono plik”. Dlatego https://example.com, https://www.example.com, https://blog.example.com i http://example.com wymagają osobnych plików.
  • Google obsługuje protokoły HTTP, HTTPS i FTP.

Podstawowe nieporozumienie: pobieranie a indeksowanie

Najważniejsza zasada brzmi: robots.txt kontroluje pobieranie, nie indeksowanie. Zablokowanie adresu URL nie usuwa go z Google. Evidence for this claim A robots.txt rule controls crawling rather than guaranteeing removal from Google Search; a URL can still appear when Google cannot crawl it. Scope: Google Search crawler behavior. Other crawlers can interpret robots.txt differently. Confidence: high · Verified: Google: Introduction to robots.txt

Dokumentacja Google mówi wprost, że robots.txt “is not a mechanism for keeping a web page out of Google. To keep a web page out of Google, block indexing with noindex or password-protect the page.” (tłumaczenie) „nie służy do usuwania strony z Google; użyj noindex lub ochrony hasłem”. Google dodaje: “While Google won’t crawl or index the content blocked by a robots.txt file, we might still find and index a disallowed URL if it is linked from other places on the web.” (tłumaczenie) „Google może znaleźć i zindeksować niedozwolony adres, jeśli prowadzą do niego odnośniki”. Wtedy “its URL can still appear in search results, but the search result won’t have a description.” (tłumaczenie) „adres może pojawić się w wynikach, ale bez opisu”. Evidence for this claim Robots.txt controls crawler access, not index eligibility; Google may still index a disallowed URL discovered through links, typically without a content snippet. Scope: web crawling Confidence: high · Verified: Robots.txt Introduction and Guide

Specyfikacja ujmuje to podobnie: “Google can’t index the content of pages which are disallowed for crawling, but it may still index the URL and show it in search results without a snippet.” (tłumaczenie) „Google nie zindeksuje treści strony niedozwolonej do pobrania, ale może zindeksować jej adres i wyświetlić go bez fragmentu”.

Dlaczego NIE wolno blokować strony przeznaczonej do noindex

noindex działa tylko wtedy, gdy Google może pobrać stronę i odczytać dyrektywę; dokumentacja Google opisuje tę zależność tak: “For the noindex rule to be effective, the page or resource must not be blocked by a robots.txt file, and it has to be otherwise accessible to the crawler. If the page is blocked by a robots.txt file or the crawler can’t access the page, the crawler will never see the noindex rule, and the page can still appear in search results, for example if other pages link to it.” (tłumaczenie) „Reguła noindex wymaga, aby robots.txt nie blokował zasobu i aby robot mógł go odczytać; inaczej adres może nadal występować w wynikach”. Evidence for this claim Google must be able to crawl a URL to see a noindex rule; blocking the URL in robots.txt can prevent the rule from being observed. Scope: Google Search indexing controls for HTML meta robots and X-Robots-Tag rules. Confidence: high · Verified: Google: Block indexing with noindex

Jeśli chcesz usunąć stronę z indeksu, zapamiętaj radę Johna Muellera: nie blokuj Google w robots.txt; użyj noindex.

Dowód z praktyki: zablokowałem dwie wysoko notowane strony

W eksperymencie z dwiema wysoko notowanymi stronami Ahrefs celowo zablokowałem je w robots.txt. Pozostały w indeksie i nadal zajmowały pozycje, lecz Google nie mógł odświeżać ich treści. “We lost a position here or there and all of the featured snippets for the pages.” (tłumaczenie) „Utraciliśmy część pozycji oraz wszystkie wyróżnione fragmenty tych stron”. Ruch spadł mniej, niż oczekiwałem: “Both pages lost some traffic. But it didn’t result in much change to our traffic estimate like I was expecting.” (tłumaczenie) „Ruch obu stron zmalał, lecz szacunek ruchu zmienił się mniej, niż zakładałem”.

Wniosek: “Accidentally blocking pages (that Google already ranks) from being crawled using robots.txt probably isn’t going to have much impact on your rankings, and they will likely still show in the search results.” (tłumaczenie) „Przypadkowa blokada stron, które już zajmują pozycje, prawdopodobnie nie wpłynie mocno na ranking i nadal będą widoczne”. Krócej: “Don’t block pages you want indexed. It hurts. Not as bad as you might think it does—but it still hurts.” (tłumaczenie) „Nie blokuj stron, które chcesz indeksować. To szkodzi, choć mniej, niż można sądzić”.

Ostrzeżenie Search Console “Indexed, though blocked by robots.txt” (tłumaczenie) „Zindeksowano mimo blokady w robots.txt” dla koszyka, filtrów lub adresów z parametrami zwykle nie jest problemem. Ma znaczenie dopiero wtedy, gdy zablokowany adres powinien być pobierany i indeksowany.

Składnia — pełne odniesienie

Robots.txt składa się z grup. Każda zaczyna się od co najmniej jednego wiersza User-agent, który wskazuje roboty, a następnie zawiera ich reguły.

User-agent: *
Disallow: /cart/
Disallow: /search
Allow: /search/help

User-agent: Googlebot
Disallow: /no-google/

Sitemap: https://example.com/sitemap.xml

User-agent i grupy. Robot stosuje dokładnie jedną grupę — najbardziej szczegółową, która pasuje do jego nazwy — i ignoruje pozostałe. “Google’s crawlers determine the correct group of rules by finding in the robots.txt file the group with the most specific user agent that matches the crawler’s user agent. Other groups are ignored.” (tłumaczenie) „Roboty Google wybierają grupę z najbardziej szczegółowym pasującym user-agentem; pozostałe grupy są ignorowane”. “Only one group is valid for a particular crawler.” (tłumaczenie) „Dla danego robota obowiązuje tylko jedna grupa”. Bing działa tak samo.

Grupa szczegółowa nie dziedziczy reguł z User-agent: *. Google stwierdza: “user agent specific groups and global groups (*) are not combined.” (tłumaczenie) „Grupy dla konkretnych robotów nie są łączone z grupą globalną”. Grupa User-agent: googlebot-news musi więc powtarzać wszystkie potrzebne reguły. Evidence for this claim For Google's crawlers, the most specific matching user-agent group applies; rules from that specific group are not combined with the global asterisk group, although multiple matching specific groups are merged internally. Scope: robots.txt parsing and fetching Confidence: high · Verified: How Google Interprets the robots.txt Specification

Disallow i Allow. Disallow wskazuje ścieżki, których robot nie może pobierać; Allow tworzy wyjątki. “specifies paths that must not be accessed by the crawlers identified by the user-agent line the disallow rule is grouped with.” (tłumaczenie) „określa ścieżki niedostępne dla robotów wskazanych przez user-agent”. Allow “specifies paths that may be accessed by the designated crawlers. When no path is specified, the rule is ignored.” (tłumaczenie) „określa dozwolone ścieżki; reguła bez ścieżki jest ignorowana”.

Reguła dopasowania. Przy konflikcie wygrywa reguła najbardziej szczegółowa, czyli z najdłuższą ścieżką. “When matching robots.txt rules to URLs, crawlers use the most specific rule based on the length of the rule path. In case of conflicting rules, including those with wildcards, Google uses the least restrictive rule.” (tłumaczenie) „Roboty wybierają najdłuższą ścieżkę, a przy konflikcie regułę najmniej restrykcyjną”. Przy remisie Allow wygrywa z Disallow. RFC 9309: “The following example shows that in the case of two rules, the longest one is used for matching.” (tłumaczenie) „Przy dwóch regułach do dopasowania używa się dłuższej”. Evidence for this claim Google resolves matching robots.txt rules by path specificity and uses the least restrictive rule when equally specific rules conflict. Scope: Google crawler interpretation of robots.txt rules; other crawlers may implement different extensions. Confidence: high · Verified: Google: Robots.txt interpretation

Przykład:

User-agent: *
Allow: /folder/page
Disallow: /folder/

Adres /folder/page pasuje do obu reguł. Allow: /folder/page ma 12 znaków, a Disallow: /folder/ osiem, więc dłuższa i bardziej szczegółowa reguła Allow wygrywa; stronę można pobrać.

Symbole * i $. * designates 0 or more instances of any valid character. $ designates the end of the URL.” (tłumaczenie) „Gwiazdka zastępuje dowolną liczbę znaków, a znak dolara wskazuje koniec adresu”. Disallow: /*.pdf$ blokuje adresy zakończone .pdf, a Disallow: /*? adresy z zapytaniem. Disallow: /fish pasuje do /fish, /fish.html i /fish/salmon.html, ale nie do /Fish ani /catfish.

Wielkość liter. Nazwy pól i robotów jej nie rozróżniają, ale wartości ścieżek tak. “Both the user-agent field name and its value are case-insensitive,” (tłumaczenie) „Nazwa i wartość pola user-agent nie rozróżniają wielkości liter”, natomiast “The field name (disallow) is case-insensitive, but its value is case-sensitive,” (tłumaczenie) „nazwa pola disallow nie rozróżnia wielkości liter, ale jego wartość tak” i “The path value must start with / to designate the root and the value is case-sensitive.” (tłumaczenie) „Ścieżka musi zaczynać się od / i rozróżnia wielkość liter”. Disallow: /Folder/ nie blokuje /folder/.

Sitemap. Dyrektywa Sitemap: przyjmuje pełny bezwzględny adres URL i nie należy do żadnej grupy, więc może znajdować się w dowolnym miejscu pliku.

Komentarze. Wszystko po # jest ignorowane: “To include comments, precede your comment with the # character.” (tłumaczenie) „Komentarz należy poprzedzić znakiem #”.

noindex, nofollow i crawl-delay NIE są dyrektywami robots.txt

To trwały mit. Od 1 września 2019 r. Google nie obsługuje nieudokumentowanych reguł noindex, nofollow ani crawl-delay. Lista obsługiwanych pól obejmuje user-agent, allow, disallow i sitemap; “other fields such as crawl-delay aren’t supported.” (tłumaczenie) „inne pola, takie jak crawl-delay, nie są obsługiwane”.

Zamiast noindex w robots.txt użyj metatagu noindex lub nagłówka X-Robots-Tag, kodu 404/410, ochrony hasłem, Disallow albo narzędzia do usuwania w Search Console — zależnie od celu.

Jak Google obsługuje robots.txt

  • Limit: 500 KiB. “Google enforces a robots.txt file size limit of 500 kibibytes (KiB). Content which is after the maximum file size is ignored.” (tłumaczenie) „Google ogranicza plik do 500 KiB; dalsza treść jest ignorowana”. RFC 9309 wymaga co najmniej tego limitu.
  • Pamięć podręczna: około 24 godzin. “Google generally caches the contents of robots.txt file for up to 24 hours, but may cache it longer in situations where refreshing the cached version isn’t possible.” (tłumaczenie) „Google zwykle przechowuje robots.txt do 24 godzin, czasem dłużej, gdy odświeżenie jest niemożliwe”. Evidence for this claim Google generally caches robots.txt for up to 24 hours and changes crawling behavior according to the HTTP status returned for the file. Scope: Google crawler handling of robots.txt fetches, including documented 4xx, 5xx, and redirect behavior. Confidence: high · Verified: Google: Robots.txt file handling
  • Kody stanu wpływają na całą witrynę. Odpowiedzi 4xx poza 429 oznaczają brak ograniczeń: “Google’s crawlers treat all 4xx errors, except 429, as if a valid robots.txt file didn’t exist. This means that Google assumes that there are no crawl restrictions.” (tłumaczenie) „Roboty traktują 4xx poza 429 jak brak pliku i zakładają brak ograniczeń”. Nie używaj odpowiedzi 401 ani 403 do ograniczania pobierania. 5xx lub niedostępność są niebezpieczne: “For the first 12 hours, Google stops crawling the site but keeps trying to fetch the robots.txt file. If Google can’t fetch a new version, for the next 30 days Google will use the last good version, while still trying to fetch a new version.” (tłumaczenie) „Przez pierwsze 12 godzin Google wstrzymuje pobieranie, a potem do 30 dni używa ostatniej poprawnej wersji, nadal próbując pobrać nową”. Po 30 dniach: “If the errors are still not fixed after 30 days: If the site is generally available to Google, Google will behave as if there is no robots.txt file (but still keep checking for a new version).” (tłumaczenie) „Jeśli błędy trwają, a witryna jest dostępna, Google zachowuje się jak przy braku robots.txt i nadal szuka nowej wersji”. Przekierowania 3xx są śledzone przez co najmniej pięć etapów, a potem traktowane jak 404.

robots.txt w Bing, Yandex i innych systemach

Grupowanie i składnia są podobne, ale liczą się trzy różnice:

  • crawl-delay. Google ją ignoruje, Bing nadal obsługuje, a Yandex wycofał w 2018 r. Yandex: “From February 22, 2018, Yandex doesn’t take into account the Crawl-delay directive,” (tłumaczenie) „Od 22 lutego 2018 r. Yandex nie uwzględnia crawl-delay” i odsyła do ustawienia szybkości pobierania w Yandex Webmaster. Bing: “The robots.txt file is the only valid place to set a crawl-delay directive for MSNBot,” (tłumaczenie) „robots.txt jest jedynym prawidłowym miejscem ustawienia crawl-delay dla MSNBot” i dyrektywa “accepts only positive, whole numbers as values… the higher the value, the more throttled down the crawl rate will be.” (tłumaczenie) „przyjmuje dodatnie liczby całkowite; wyższa wartość silniej ogranicza tempo”. Bing traktuje wartość jako względne ograniczenie, nie dosłowną liczbę sekund.
  • Pułapka grupy bingbot. Zgodnie z regułą “only one group per crawler” (tłumaczenie) „tylko jedna grupa dla danego robota”, po dodaniu User-agent: bingbot Bing stosuje tylko tę grupę i ignoruje reguły User-agent: * poza crawl-delay, więc potrzebne dyrektywy trzeba powtórzyć.
  • Pamięć podręczna Amazon. Roboty Amazon mogą używać kopii z ostatnich 30 dni. Gdy pliku nie da się pobrać, zachowują się jak przy jego braku. Tester nie potwierdzi, którą kopię widział Amazon. Evidence for this claim Amazon says its crawlers may use a robots.txt copy cached within the previous 30 days and behave as though the file does not exist when they cannot fetch it. Scope: Amazon crawler behavior only; a checker result cannot establish which cached copy Amazon used or whether Amazon observed the same fetch failure. Confidence: high · Verified: Amazon: Amazonbot

Zarządzanie robotami AI za pomocą robots.txt

Robots.txt jest obecnie głównym mechanizmem zarządzania robotami AI. Używają tej samej składni grup i user-agentów, ale każdy robot ma oddzielny identyfikator, więc zablokowanie jednego nie blokuje pozostałych.

  • OpenAI używa niezależnych robotów. GPTBot pobiera treści do trenowania modeli; OAI-SearchBot zasila funkcje wyszukiwania ChatGPT; OAI-AdsBot sprawdza bezpieczeństwo stron reklamowych i nie służy do trenowania. User-agent: GPTBot z Disallow: / blokuje trening, ale nie roboty wyszukiwania ani reklam. ChatGPT-User uruchamia się na żądanie użytkownika w ChatGPT lub Custom GPT, a OpenAI zaznacza, że “robots.txt rules may not apply” (tłumaczenie) „reguły robots.txt mogą nie mieć zastosowania”. Zmiana dostępu OAI-SearchBot może docierać do systemów wyszukiwania około 24 godzin.
  • Google-Extended steruje użyciem w Gemini i Vertex, niezależnie od Googlebota.
  • Inne identyfikatory to CCBot, ClaudeBot, PerplexityBot i Bytespider.

Najważniejsze zastrzeżenie: zgodność jest dobrowolna. Robots.txt prosi, ale nie wymusza. Rzetelne roboty go respektują, lecz scraper może go zignorować. Rzeczywista ochrona wymaga uwierzytelniania lub blokowania dostępu.

Typowe błędy i ich poprawki

Plik zwraca 200, ale nie jest użytecznym robots.txt. Sam status nie wystarcza. Sprawdź Content-Type i pierwsze bajty, bo CDN może zwrócić HTML błędu z kodem 200. Google oczekuje zwykłego tekstu UTF-8 i może ignorować nieprawidłowe znaki. Pojedynczy BOM UTF-8 na początku jest tolerowany, ale drugi BOM, BOM w środku, UTF-16, bajty NUL lub znaki sterujące mogą zmienić token albo unieważnić wiersz. Raportuj pozycję bajtu i wiersz bez cichej normalizacji. Przed obliczeniem reguł zastosuj limit parsowania 500 KiB Google, a odrzucony fragment nadal pokaż w raporcie.

  • Blokowanie strony przeznaczonej do noindex. Google nie pobierze jej i nie zobaczy noindex; usuń blokadę.
  • Używanie robots.txt do usuwania z indeksu. Do tego służy noindex.
  • Blokowanie CSS/JS potrzebnych do renderowania. Google musi pobrać te zasoby; przykładowy plik Google jawnie zezwala na .css i .js.
  • Ukrywanie poufnych danych. “The Robots Exclusion Protocol is not a substitute for valid content security measures. Listing paths in the robots.txt file exposes them publicly and thus makes the paths discoverable.” (tłumaczenie) „Robots Exclusion Protocol nie zastępuje zabezpieczeń; wpisanie ścieżek ujawnia je publicznie”. Wpisanie /secret-admin/ publicznie ujawnia ścieżkę; użyj uwierzytelniania.
  • Przypadkowe Disallow: /. Blokuje całą witrynę dla wskazanego robota.
  • Ignorowanie statusu /robots.txt. Błąd 5xx może wstrzymać pobieranie całej witryny.

Szerszy kontekst — odkrywanie, harmonogram pobierania, renderowanie oraz różnica między pobieraniem a indeksowaniem — opisuje centrum dotyczące crawlowania. Osobne materiały rozwijają budżet indeksowania i obsługę map witryn przez Google.

Who's been ignoring my robots.txt?

This is live data from this site, not an illustration. My robots.txt disallows /api/trap/, and the only link to it is invisible to humans — so a compliant crawler will never request it. Every user-agent below fetched it anyway. (Humans poking at it with curl show up too; the user-agent usually gives them away.)

Loading trap log…

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.