Türkçe çeviri: Robots.txt

ne robots.txt aslında yapar — o controls tarama, değil dizine ekleme — plus exact syntax, Google'ın nasıl ele aldığı o altında hood, ve mistakes şu break siteler.

İlk yayın tarihi: 23 Haz 2026 · Son güncelleme: 8 Ağu 2026 · Advanced
Diller
Bu sayfada 1 kanıt sinyali

Robots.txt dır bir plain-text file at root of her ana makine şu söyler crawlers hangi URLs onlar -ebilir ve -ebilir değil istek. bir thing -e al yapğru: o controls tarama, değil dizine ekleme. bir disallowed URL -ebilir hâlâ olmak dizine eklenmiş olmadan bir snippet eğer o's linked -den elsewhere — -e koru bir sayfa out of dizin siz kullan noindex, ve sayfa -meli değil olmak blocked in robots.txt veya Google never sees noindex. Google supports yalnızca kullanıcı-agent, izin ver, disallow, ve sitemap (noindex, nofollow, ve tarama-delay idi dropped Sept 1, 2019). o lives at /robots.txt, dır scoped -e bir ana makine+protokol+port, caps at 500 KiB, caches ~24h, ve bir 4xx anlamına gelir no restrictions -iken bir 5xx -ebilir stall tarama site-wide. yapmayın block render-critical CSS/JS, ve yapmayın rely on o -e hide anything — file dır kamuya birçık.

TL;DR — Robots.txt dır bir plain-text file at root of her ana makine (/robots.txt, lowercase) şu implements Robots Exclusion protokol (RFC 9309). o controls tarama, değil dizine ekleme — bir disallowed URL -ebilir hâlâ olmak dizine eklenmiş olmadan bir snippet eğer linked elsewhere; -e deindex kullan noindex on bir sayfa bu değil blocked. Google supports yalnızca user-agent, allow, disallow, ve sitemap; noindex/nofollow/crawl-delay idi dropped Sept 1, 2019. Scope dır bir ana makine+protokol+port. Matching kullanır en çok-specific (longest) kural, least-restrictive on ties; * ve $ dır wildcards; paths dır durum-sensitive. Google caps file at 500 KiB, caches ~24h, treats 4xx (except 429) olarak no-restrictions, ve on bir 5xx stalls tarama bençin ~12h o hâlde falls back -e son good kopya bençin ~30 days. yapmayın block render-critical CSS/JS, ve yapmayın ele al o olarak access control — file dır kamuya birçık.

ne o dır ve nerede o lives

Robots.txt implements Robots Exclusion protokol, oluşturuldu tarafından Martijn Koster in 1994 ve finally standardized in 2022 olarak RFC 9309 — co-authored tarafından Google’s Gary Illyes, Henner Zeller, Lizzi Sassman, ve Koster himself. standard’s kendi wording: “bu belge specifies ve extends ‘Robots Exclusion protokol’ yöntem originally defined tarafından Martijn Koster in 1994 bençin service owners -e control nasıl bençerik sunulan tarafından onların services -ebilir olmak accessed, eğer at tümü, tarafından automatic clients known olarak crawlers.”

bir few facts şu catch kişiler out:

  • o -meli olmak at root, lowercase. RFC 9309 dır explicit: ” kurallar -meli olmak accessible in bir file named ‘/robots.txt’ (tümü lowercase) in top-level path of service.” Google adds şu URL itself dır durum-sensitive, like herhangi bir URL.
  • Scope dır bir ana makine + protokol + port. Google: ” kurallar listed in robots.txt file uygula yalnızca -e ana makine, protokol, ve port number nerede robots.txt file dır hosted.” bu nedenle https://example.com, https://www.example.com, https://blog.example.com, ve http://example.com her ihtiyaç duy onların kendi file. Subdomains ve protocols yapmayın share bir.
  • Supported protocols bençin Google dır HTTP, HTTPS, ve FTP.

misconception şu defines bu topic: tarama vs dizine ekleme

-erseniz take bir thing -den bu sayfa, take bu: robots.txt controls tarama, değil dizine ekleme. Blocking bir URL değildir aynı olarak removing o -den Google. Evidence for this claim A robots.txt rule controls crawling rather than guaranteeing removal from Google Search; a URL can still appear when Google cannot crawl it. Scope: Google Search crawler behavior. Other crawlers can interpret robots.txt differently. Confidence: high · Verified: Google: Introduction to robots.txt

Google’ın kendi intro doc söyler o plainly: robots.txt “değildir bir mechanism bençin tutma bir web sayfası out of Google. -e koru bir web sayfası out of Google, block dizine ekleme ile noindex veya password-protect sayfa.” ve on ne aslında olur -e bir blocked URL: “-iken Google won’t tarama veya dizin bençerik blocked tarafından bir robots.txt file, biz -ebilir hâlâ bul ve dizin bir disallowed URL eğer o dır linked -den diğer places on web.” The result is the familiar snippet-less listing: “onun URL -ebilir hâlâ görün in arama sonuçları, ama arama sonuç won’t sahip bir description.”

Evidence for this claim Robots.txt controls crawler access, not index eligibility; Google may still index a disallowed URL discovered through links, typically without a content snippet. Scope: web crawling Confidence: high · Verified: Robots.txt Introduction and Guide

spec restates aynı nuance bençin disallow kural itself: “Google -ebilir’t dizin bençerik of sayfalar hangi dır disallowed bençin tarama, ama o -ebilir hâlâ dizin URL ve göster o in arama sonuçları olmadan bir snippet.”

neden siz -meli değil block bir sayfa siz iste -e noindex

bu trap şu quietly breaks deindexing efforts. bir noindex yalnızca çalışır eğer Google -ebilir tarama sayfa -e okuyun o. Google’s block-dizine ekleme doc spells out dependency: “bençin noindex kural -e olmak effective, sayfa veya kaynak -meli değil olmak blocked tarafından bir robots.txt file, ve o sahiptir -e olmak otherwise accessible -e crawler. eğer sayfa dır blocked tarafından bir robots.txt file veya crawler -ebilir’t access sayfa, crawler -ecek never see noindex kural, ve sayfa -ebilir hâlâ görün in arama sonuçları, örneğin eğer diğer sayfalar bağlantı -e o.” Evidence for this claim Google must be able to crawl a URL to see a noindex rule; blocking the URL in robots.txt can prevent the rule from being observed. Scope: Google Search indexing controls for HTML meta robots and X-Robots-Tag rules. Confidence: high · Verified: Google: Block indexing with noindex

bu nedenle eğer sizin goal dır -e al bir sayfa out of dizin, John Mueller’s rehberlik dır cleanest way -e remember o: -dığınızda iste -e unindex sayfalar, yapmalısınız değil block Google ile robots.txt, ama rather kullan noindex.

lived proof: ben blocked two of bizim kendi high-sıralama sayfalar

ben yapmayın sahip -e argue bu -den theory. In benim experiment blocking two high-sıralama Ahrefs sayfalar, ben deliberately blocked them in robots.txt ve tracked ne happened. sayfalar stayed dizine eklenmiş ve kept sıralama — onlar didn’t vanish. ne biz lost idi freshness Google alır -den re-tarama: “biz lost bir position burada veya orada ve tümü of featured snippets bençin sayfalar.” trafik dropped, ama -den az ben beklenen: “her ikisi sayfalar lost bazı trafik. ama o didn’t sonuç in much change -e bizim trafik estimate like ben idi expecting.”

benim takeaway -den data: “Accidentally blocking sayfalar (şu Google zaten ranks) -den olma tarandı kullanarak robots.txt probably değildir going -e sahip much impact on sizin sıralamalar, ve onlar -ecek muhtemel hâlâ göster in arama sonuçları.” ve blunt sürüm: “yapmayın block sayfalar siz iste dizine eklenmiş. o hurts. değil olarak bad olarak siz -ebilir think o yapar—ama o hâlâ hurts.”

flip side dır reassurance: ne zaman arama Console flags “dizine eklenmiş, gerçben blocked tarafından robots.txt” bençin bir utility URL — cart, filter, parameter junk — o’s genellikle bir non-sorun. olarak Mueller put o hakkında ekle—e-cart URLs, blocking them dır fine, ve hatta eğer onlar al “indexed,” o’s unlikely onlar’ll olmak gösterilen in arama unless someone runs bir çok specific sorgu bençin şunlar URLs, hangi gerçek kullanıcılar yapmayın yap. Distinguish scary-sounding warning -den bir gerçek sorun: o yalnızca önem taşır eğer blocked URL dır bir sayfa siz aslında wanted tarandı ve dizine eklenmiş.

syntax ( reference)

bir robots.txt dır bir ayarla of gruplar. her grup starts ile bir veya daha User-agent lines naming hangi crawler(s) o uygulanır -e, izlenen tarafından kurallar bençin them.

User-agent: *
Disallow: /cart/
Disallow: /search
Allow: /search/help

User-agent: Googlebot
Disallow: /no-google/

Sitemap: https://example.com/sitemap.xml

kullanıcı-agent ve gruplar. bir crawler obeys tam olarak bir grup — bir ile en çok specific kullanıcı-agent şu matches o — ve ignores rest. Google: “Google’s crawlers determine correct grup of kurallar tarafından bulma in robots.txt file grup ile en çok specific kullanıcı agent şu matches crawler’s kullanıcı agent. diğer gruplar dır ignored.” And: “yalnızca bir grup dır valid bençin bir particular crawler.” (Bing behaves aynı way — daha on şu below.)

şu ayrıca anlamına gelir bir specific grup yapmaz al topped up ile wildcard grup’s kurallar — o’s kullanılan kendi başına, değil merged ile User-agent: *. Google’s spec dır explicit şu “kullanıcı agent specific gruplar ve global gruplar (*) değildir combined.” bu nedenle -erseniz yaz bir User-agent: googlebot-news grup, o sahiptir -e olmak self-contained: anything siz hâlâ iste o -e obey -den * grup sahiptir -e olmak repeated bençinde o, veya Googlebot-News basitçe won’t see şunlar kurallar at tümü.

Evidence for this claim For Google's crawlers, the most specific matching user-agent group applies; rules from that specific group are not combined with the global asterisk group, although multiple matching specific groups are merged internally. Scope: robots.txt parsing and fetching Confidence: high · Verified: How Google Interprets the robots.txt Specification

Disallow ve izin ver. Disallow listeler paths bir crawler -meli değil istek; Allow carves exceptions back out. disallow kural “specifies paths şu -meli değil olmak accessed tarafından crawlers identified tarafından kullanıcı-agent line disallow kural dır grouped ile.” The allow rule “specifies paths şu -ebilir olmak accessed tarafından designated crawlers. ne zaman no path dır specified, kural dır ignored.”

** matching kural (en çok guides al bu yanlış).** ne zaman two kurallar conflict, en çok specific bir wins, ve “most specific” anlamına gelir longest path: “ne zaman matching robots.txt kurallar -e URLs, crawlers kullan en çok specific kural based on length of kural path. In durum of conflicting kurallar, dahil şunlar ile wildcards, Google kullanır least restrictive kural.” bu nedenle on bir genuine tie, least restrictive kural wins — Allow beats Disallow. RFC 9309 frames o olarak “Longest Match”: ” following örnek shows şu in durum of two kurallar, longest bir dır kullanılan bençin matching.” Evidence for this claim Google resolves matching robots.txt rules by path specificity and uses the least restrictive rule when equally specific rules conflict. Scope: Google crawler interpretation of robots.txt rules; other crawlers may implement different extensions. Confidence: high · Verified: Google: Robots.txt interpretation

Worked örnek:

User-agent: *
Allow: /folder/page
Disallow: /folder/

URL /folder/page matches her ikisi kurallar. Allow: /folder/page (12 chars) dır longer -den Disallow: /folder/ (8 chars), bu nedenle longer, daha specific izin ver wins ve sayfa dır crawlable.

Wildcards * ve $. Google: * designates 0 veya daha instances of herhangi bir valid character. $ designates end of URL.” bu nedenle Disallow: /*.pdf$ blocks her URL ending in .pdf, ve Disallow: /*? blocks her URL containing bir sorgu string. Matching dır prefix-based: Disallow: /fish matches /fish, /fish.html, ve /fish/salmon.html, ama değil /Fish (durum-sensitive) veya /catfish (o’s bir prefix, değil bir substring).

durum sensitivity ( subtle bir). Field ve kullanıcı-agent names dır durum-insensitive; path values dır durum-sensitive. Google: “her ikisi user-agent field name ve onun değer dır durum-insensitive,” but ” field name (disallow) dır durum-insensitive, ama onun değer dır durum-sensitive,” and ” path değer -meli başla ile / -e designate root ve değer dır durum-sensitive.” bu nedenle Disallow: /Folder/ yapmaz block /folder/.

Sitemap. Sitemap: directive takes bir full absolute URL ve dır independent of gruplar — o -ebilir sit anywhere in file.

Comments. Anything sonra # dır ignored: “-e bençer comments, precede sizin comment ile # character.”

noindex, nofollow, ve tarama-delay değildir robots.txt directives

bu bir persistent myth. olarak of September 1, 2019, Google retired support bençin unsupported, undocumented kurallar — dahil noindex, nofollow, ve crawl-delay. Google’s announcement focused on kurallar unsupported tarafından internet draft, such olarak tarama-delay, nofollow, ve noindex, noting onlar idi never documented tarafından Google, ve said Google idi retiring tümü code şu handles unsupported ve unpublished kurallar (such olarak noindex) on şu date. supported field liste dır kısa, ve spec çbirğrılar out exclusion yapğrudan: Google supports user-agent, allow, disallow, ve sitemap, ve “diğer fields such olarak crawl-delay değildir supported.”

-erseniz relied on noindex in robots.txt, alternatives dır bir noindex meta tag veya X-Robots-Tag header, 404/410 status codes, password protection, bir Disallow, veya arama Console kaldırma araç.

Google’ın nasıl ele aldığı robots.txt altında hood

  • Size limit: 500 KiB. “Google enforces bir robots.txt file size limit of 500 kibibytes (KiB). bençerik bu da sonra maximum file size dır ignored.” RFC 9309 aligns: “The parsing limit MUST be at least 500 kibibytes [KiB].”
  • Caching: ~24 hours. “Google generally caches contents of robots.txt file bençin up -e 24 hours, ama -ebilir cache o longer in situations nerede refreshing cached sürüm değildir olası.” bu nedenle bir change değildir necessarily seçilen up instantly. Evidence for this claim Google generally caches robots.txt for up to 24 hours and changes crawling behavior according to the HTTP status returned for the file. Scope: Google crawler handling of robots.txt fetches, including documented 4xx, 5xx, and redirect behavior. Confidence: high · Verified: Google: Robots.txt file handling
  • Status codes önem taşır site-wide. bu part en çok guides skip:
    • 4xx (except 429) → no restrictions. “Google’s crawlers ele al tümü 4xx errors, except 429, olarak eğer bir valid robots.txt file didn’t var ol. bu anlamına gelir şu Google assumes şu vardır no tarama restrictions.” bir 404 on /robots.txt anlamına gelir “crawl everything.” (yapmayın kullan 401/403 -e throttle tarama.)
    • 5xx / unreachable → dangerous. “bençin ilk 12 hours, Google stops tarama site ama keeps trying -e fetch robots.txt file. eğer Google -ebilir’t fetch bir yeni sürüm, bençin sonraki 30 days Google -ecek kullan son good sürüm, -iken hâlâ trying -e fetch bir yeni sürüm.” bu nedenle bir server error on /robots.txt -ebilir effectively disallow sizin whole site bençin ilk ~12 hours, o hâlde çalıştır on son cached kopya bençin ~30 days. bir persistently erroring robots.txt dır bir site-wide tarama risk. ve eğer o’s hâlâ broken sonra şunlar 30 days: “eğer errors dır hâlâ değil düzeltilmiş sonra 30 days: eğer site dır generally mevcut -e Google, Google -ecek behave olarak eğer yoktur robots.txt file (ama hâlâ koru checking bençin bir yeni sürüm).” In diğer words, bir robots.txt şu never recovers yapmaz stay disallowed forever — Google eventually falls back -e tarama ile no restrictions, aynı olarak bir 404.
    • 3xx → Google follows at least five yönlendirme hops, o hâlde treats o olarak bir 404.

robots.txt in Bing, Yandex, ve beyond

grouping ve syntax dır essentially shared, ama two divergences önem taşır:

  • tarama-delay. Google ignores o, Bing hâlâ honors o, ve Yandex dropped o in 2018 — Yandex’s kendi dokümantasyon states şu “-den February 22, 2018, Yandex yapmaz take -e account tarama-delay directive,” pointing siz -e site tarama rate ayarlama in Yandex Webmaster instead. Bing dır explicit şu ” robots.txt file dır yalnızca valid place -e ayarla bir tarama-delay directive bençin MSNBot,” and that the directive “accepts yalnızca positive, whole numbers olarak values… higher değer, daha throttled down tarama rate -ecek olmak.” Note Bing treats değer olarak bir relative throttle, değil literally N seconds.
  • ** bingbot-section gotcha.** sadece like Google’s “only one group per crawler” kural, -erseniz oluştur bir User-agent: bingbot section, Bing uygulanır yalnızca şu section ve ignores User-agent: * defaults (tarama-delay excepted). bu nedenle bir bingbot-specific grup -meli repeat her directive siz hâlâ iste enforced.
  • Amazon’s cache ve başarısızlık behavior. Amazon söyler onun crawlers -ebilir kullan bir robots.txt kopya cached bençinde previous 30 days. eğer onlar cannot fetch file, onlar behave olarak gerçben o yapmaz var ol. bir checker -ebilir rapor kopya o fetched, ama o cannot prove hangi cached sürüm Amazon kullanılan—veya şu Amazon observed aynı başarısızlık olarak checker. Evidence for this claim Amazon says its crawlers may use a robots.txt copy cached within the previous 30 days and behave as though the file does not exist when they cannot fetch it. Scope: Amazon crawler behavior only; a checker result cannot establish which cached copy Amazon used or whether Amazon observed the same fetch failure. Confidence: high · Verified: Amazon: Amazonbot

Managing AI crawlers ile robots.txt

Robots.txt dır currently main lever bençin managing AI crawlers, ve onlar obey aynı grup/kullanıcı-agent syntax. catch: bunlar dır separate tokens, bu nedenle blocking bir yapmaz block others.

  • OpenAI runs several distinct bots, ve controls her biri bençin dır independent — allowing bir yapmaz izin ver others, ve blocking bir yapmaz block others. GPTBot tarar bençerik bençin training OpenAI’s models; OAI-SearchBot surfaces siteler in ChatGPT’s arama features; OAI-AdsBot kontroller safety of sayfalar submitted olarak ads (onun data değildir kullanılan bençin training). Block training ile User-agent: GPTBot / Disallow: / — şu alone won’t durdur arama veya ads bots. ChatGPT-User dır farklı yeniden: o fires bençin actions bir kişben triggers bençinde ChatGPT veya bir Custom GPT, değil automatic tarama, ve OpenAI söyler “robots.txt rules may not apply” -e o — bu nedenle yapmayın count on bir Disallow -e koru o out. -erseniz yap change ne OAI-SearchBot -ebilir tarama, OpenAI notes o -ebilir take hakkında 24 hours bençin update -e ulaş onların arama systems.
  • Google-Extended controls Gemini/Vertex training ve dır separate -den Googlebot.
  • Others worth naming: CCBot (yaygın tarama), ClaudeBot (Anthropic), PerplexityBot, ve Bytespider.

hard caveat: compliance dır voluntary. Robots.txt istekler; o yapmaz enforce. Well-behaved crawlers obey o; scrapers -ebilir ve yap ignore o. -erseniz truly ihtiyaç duy -e koru something away -den bir bot, şu’s bir authentication/blocking sorun, değil bir robots.txt bir.

yaygın mistakes (ve düzeltmeler)

** file dır 200, ama o değildir aslında bir usable robots file.** Status alone dır değil enough. Capture response Content-Type ve ilk bytes: bir CDN/custom-error template -ebilir döndür HTML at /robots.txt ile 200, hangi -meli olmak bir warning rather -den bir “allow all” geç. Google belgeler robots.txt olarak UTF-8 plain text ve -ebilir ignore invalid characters. bir tek UTF-8 BOM at beginning dır tolerated, ama bir ikinci BOM, bir BOM in middle, UTF-16 bytes, NULs, veya invisible/control characters -ebilir alter ilk token veya invalidate bir line. göster byte offset ve affected line; yapmayın silently normalize file önce telling kullanıcı ne crawler received. uygula Google’s effective 500 KiB parsing limit önce calculating izin ver/disallow sonuçlar, -iken hâlâ reporting discarded tail.

  • Blocking bir sayfa siz ayrıca iste deindexed. Block + noindex anlamına gelir Google never tarar o -e see noindex. kullan noindex olmadan block.
  • kullanarak robots.txt -e deindex. yanlış araç entirely — şu’s noindex’s job.
  • Blocking render-critical CSS/JS. Google gerektirir şunlar assets -e see sayfa olarak bir kullanıcı yapar; Google’ın kendi sample robots.txt explicitly re-izin verir .css/.js bu nedenle Googlebot -ebilir tarama them.
  • Trying -e hide sensitive data. RFC 9309 dır blunt: ” Robots Exclusion protokol değildir bir substitute bençin valid bençerik security measures. Listing paths in robots.txt file exposes them publicly ve thus yapar paths discoverable.” Disallowing /secret-admin/ literally advertises o. kullan auth.
  • bir stray Disallow: /. bu blocks entire site bençin named crawler — classic staging leftover şu takes bir site out of Google.
  • Ignoring response code on /robots.txt. bir 5xx -ebilir stall tarama site-wide; ele al file’s availability olarak production-critical.

bençin broader pipeline bu sits bençinde — discovery, tarama scheduler, rendering, ve nasıl tarama differs -den dizine ekleme — see tarama hub. sibling topics (tarama bütçesi, ve Google’ın nasıl ele aldığı sitemaps) her go deeper on bir piece of bu.

Who's been ignoring my robots.txt?

This is live data from this site, not an illustration. My robots.txt disallows /api/trap/, and the only link to it is invisible to humans — so a compliant crawler will never request it. Every user-agent below fetched it anyway. (Humans poking at it with curl show up too; the user-agent usually gives them away.)

Loading trap log…

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.