Türkçe çeviri: Log File Analysis

nasıl -e okuyun sizin server's raw access logs -e see tam olarak ne Googlebot, Bingbot, ve AI crawlers aslında fetched — verifying gerçek bots, bulma tarama waste ve orphans, ve neden logs dır ground truth şu tarama araçlar ve arama Console yalnızca approximate.

İlk yayın tarihi: 22 Haz 2026 · Son güncelleme: 8 Ağu 2026 · Advanced
Diller
Bu sayfada 1 kanıt sinyali

Log file analysis dır okuma sizin server's raw access logs — unsampled, ground-truth record of her istek server aldı — -e see tam olarak hangi URLs Googlebot, Bingbot, ve AI crawlers aslında fetched, nasıl çoğu zaman, ve ile ne status code. non-negotiable ilk adım dır verifying bots dır gerçek (reverse + forward DNS, veya Google's published IP ranges), çünkü kullanıcı agents dır spoofed tümü time. o hâlde siz bak bençin tarama waste, en çok/least-tarandı URLs, status codes tarafından frequency, orphan sayfalar, ve mobile-vs-desktop split. Logs complement GSC tarama istatistikler; onlar yapmayın replace o. en çok küçük siteler yapmayın ihtiyaç duy bu — o's bir büyük-site, ecommerce, ve migration araç.

TL;DR — Logs dır unsampled ground truth bençin tarama: her istek, her bot, her status code. non-negotiable ilk adım dır verifying Googlebot/Bingbot via reverse + forward DNS (veya Google’s published IP-range JSON) — kullanıcı agents dır spoofed constantly, ve siz yap tümü tarama math on verified ayarla yalnızca. o hâlde okuyun logs bençin en çok/least-tarandı URLs ve sections, tarama frequency üzerinde time, status codes prioritized tarafından frequency, tarama waste (params, facets, site bençben arama, infinite sayfalama), orphan ve uncrawled sayfalar (cross-referenced karşı bir tarama), ve mobile-vs-desktop Googlebot split. Bing publishes no resmî IP ranges, bu nedenle DNS -e *.search.msn.com dır yöntem orada. Logs complement GSC tarama istatistikler — onlar yapmayın replace o. ve in 2026, AI bots dır now bir huge share of ne shows up.

Evidence for this claim Web-server access logs record HTTP requests and commonly include request, response-status, user-agent, and timing fields depending on configuration. Scope: Apache HTTP Server access-log behavior; other servers vary by configuration. Confidence: high · Verified: Apache HTTP Server: Log Files Evidence for this claim User-agent text alone does not authenticate Googlebot; Google recommends DNS verification or matching published IP ranges. Scope: Google crawler verification, applicable when classifying log traffic. Confidence: high · Verified: Google Search Central: Verify Googlebot

neden logs dır ground truth

vardır three ways -e “see” nasıl arama motorları tarama siz, ve onlar değildir equal:

  • bir tarama araç (Screaming Frog SEO Spider, Ahrefs site denetimi) simulates bir tarama. o söyler siz ne bir bot -ebilirdi bul, değil ne Google yaptı fetch.
  • GSC tarama istatistikler summarizes gerçek thing, ama o’s sampled, aggregated, ve capped (kabaca 1 000 rows, ~90 days, no per-URL export).
  • Server logs record gerçek thing — her istek, bençin her bot, ile exact URL, timestamp, ve status code.

Ahrefs log-file rehber ben reviewed puts o plainly: server logs dır ” en çok trustworthy kaynak of information -e understand URLs şu arama motorları sahip tarandı.” şu’s whole neden bu technique vardır. ne zaman ben iste -e know ne Googlebot aslında yaptı — değil ne o -ebilir yap, değil bir rounded özet — ben go -e logs.

bir typical log line carries IP adres, kullanıcı agent, URL path, timestamp, istek yöntem (al/POST), ve HTTP status code. Everything below dır sadece slicing şunlar fields intelligently.

-dığınızda aslında ihtiyaç duy o (ve -dığınızda yapmayın)

olmak honest ile yourself burada. Log file analysis dır bir büyük-site araç. o earns onun koru on siteler ile tens of thousands of URLs, ecommerce ve fasetli gezinme, siteler mid-migration, ve siteler stuck in Discovered – currently not indexed. olarak ben wrote in benim tarama bütçesi rehber, “en çok siteler yapmayın ihtiyaç duy -e worry hakkında tarama budget, ama vardır few durumlar nerede siz -ebilir iste -e take bir bak.” Daniel Waisberg of Google sahiptir yapılmış bir benzer benşaret et hakkında tarama istatistikler — per arama motoru Journal’s coverage, rapor değildir much of bir concern bençin siteler altında ~1 000 sayfalar.

eğer sizin few-hundred-sayfa site dır olma tarandı fine, skip bu ve go düzelt something ile daha leverage.

nasıl -e al sizin logs ( hardest part dır access)

Logs live wherever istek aslında terminated:

  • Apache & Nginx → Apache “combined” log format ( en çok yaygın).
  • Microsoft IIS → W3C format.
  • AWS ELB/ALB → ELB format.
  • CDNs (Cloudflare, Fastly, Akamai) → onların kendi log exports. bu önem taşır: at bir CDN-fronted site, bir origin-yalnızca log misses edge-cached hits, bu nedenle pull logs at layer bot aslında ulaşılan.

Aim bençin 30 days minimum, 90 ideal, -e capture tarama-frequency variation. ve plan bençin friction — getting access -e server logs dır çoğu zaman genuinely hard part (DevOps gatekeeping). hatta Googlers, in bir migration episode of arama Off Record, flagged nasıl hard log files -ebilir olmak -e obtain uygulamada. Budget time bençin istek.

Logs değildir sadece bot trafik — onlar capture her istek, dahil gerçek visitors, ve -ebilir carry sorgu-string values, session identifiers, veya diğer sensitive data alongside URL path. OWASP’s logging rehberlik dır blunt hakkında bu: authentication credentials, access tokens, ve personally identifiable information generally shouldn’t land yapğrudan in bir log; onlar -meli olmak removed, masked, veya hashed ilk. oluştur şu -e sizin access controls ve export süreç önce siz hand bir log file -e anyone bençin analysis, değil sonra.

adım 1 — Verify bots dır gerçek (yap bu önce anything else)

bu adım en çok guides wave at in bir line. yapmayın. çok söyleıda bots pretend -e olmak Googlebot -e al past firewalls (Ahrefs). kullanıcı agent dır unauthenticated text; ele al her “Googlebot” line olarak bir claim -e olmak proven.

Googlebot — iki geçerli yöntem:

  1. Reverse + forward DNS ( bidirectional kontrol et). Google’ın kendi adımlar: çalıştır bir reverse DNS lookup on IP -den sizin logs ile host command; verify domain dır googlebot.com, google.com, veya googleusercontent.com; o hâlde çalıştır bir forward DNS lookup on şu hostname ve verify o resolves back -e özgün IP. forward adım dır ne yapar bu trustworthy — bir spoofer -ebilir benşaret et reverse DNS at bir *.googlebot.com name, ama yalnızca round-trip back -e aynı IP proves o. (Commands bençin macOS/Linux ve Windows dır in Scripts tab.)
  2. Match karşı Google’s published IP ranges. Google publishes JSON files of onun crawler IPs in CIDR format — common-crawlers.json bençin Googlebot ve friends, plus special-crawlers.json, kullanıcı-triggered-fetcher files, ve bir tümü-Google goog.json. olarak ben noted in benim Googlebot rehber, Google “provided bir liste of kamuya birçık IPs -ebilirsiniz kullan -e verify istekler dır -den Google… -ebilirsiniz compare bu -e data in sizin server logs.”

Bingbot — DNS yalnızca. bu key contrast: Bing yapmaz officially publish IP ranges. Bing’s kendi wording dır şu “…like diğer arama motorları, Bing yapmaz publish bir liste of IP adresler veya ranges -den hangi biz tarama Internet,” because ” IP adresler veya ranges biz kullan -ebilir change herhangi bir time.” bu nedenle bençin Bingbot siz yap reverse + forward DNS -e bir hostname ending in *.search.msn.com (e.g. msnbot-157-55-33-18.search.msn.com), veya kullan Verify Bingbot araç. (Microsoft sahiptir since released bir bingbot IP JSON, ama onun resmî verification rehberlik hâlâ centers on DNS precisely çünkü IPs change.)

o hâlde throw out fakes. yap tümü of sizin tarama math on verified ayarla yalnızca. Unverified “Googlebot” dır almost her zaman bir scraper veya bir spoofed bot ve belongs in bir security review, değil sizin tarama-waste analysis.

adım 2 — ne -e bak bençin

Once siz’re working ile verified hits, burada’s okuyun:

  • en çok & least tarandı URLs ve sections. Rank istekler tarafından URL ve tarafından directory. bu nerede sizin tarama bütçesi dır aslında going — ve o’s genellikle surprising.
  • tarama frequency üzerinde time. Trend tarar tarafından URL/section -e catch drops (bir migration broke something) veya spikes (bir yeni section, veya bir spider trap spinning up infinite URLs).
  • Status codes bots hit, prioritized tarafından frequency. Quantify 200 vs. 301/302 (ve chains), 404, ve 5xx. bir 404 hit 5 000×/week dır bir farklı sorun -den bir 404 hit once — düzelt tarafından tarama frequency, değil tarafından mere existence.
  • tarama waste. Faceted nav, URL parametreleri, site bençben arama sonuçlar, ve infinite calendars/sayfalama -ebilir eat bir büyük share of tarama bütçesi on bad offenders. Logs göster tam olarak hangi junk patterns bots dır burning time on.
  • Orphan & uncrawled sayfalar. bu gerektirir her ikisi datasets. Cross-reference logs karşı bir site tarama: URLs in logs ama değil in tarama = orphans, eski yönlendirmeler, veya externally-linked sayfalar; URLs in tarama ama değil in logs = sayfalar Google sahiptir never fetched.
  • Mobile vs. desktop Googlebot. Split tarafından kullanıcı agent. Post mobile-ilk, o -meli olmak majority Googlebot Smartphone — bir desktop-heavy split dır worth bir bak.
  • Response time & tarama health. Rising average response time correlates ile reduced tarama. Per SEJ’s yaz-up of Waisberg’s rehberlik: “Watch out bençin bir consistent increase in average response time. Google şunu söylüyor o -ebilir değil affect tarama rate immediately, ama o’s bir good indicator şu sizin servers -ebilir değil olmak handling tümü load.”

ne logs yapmayın söyle siz

koru bunlar straight veya siz’ll üzerinde-okuyun data:

  • tarama ≠ dizin. bir URL Googlebot fetches daily -ebilir stay unindexed indefinitely. Logs prove fetching, değil dizin status — pair them ile GSC’s sayfa dizine ekleme / URL Inspection -e öğren dizin side.
  • tarama ≠ rank, ve daha tarama yapmaz yardım et. olarak ben’ve said repeatedly, “The rate of crawling isn’t going to impact your rankings.” yapmayın chase tarama volume olarak eğer o idi bir sıralama lever.
  • noindex yapmaz reduce tarama. noindex controls indexation, değil tarama — -e aslında durdur tarama siz kullan robots.txt veya bir status code.
  • tarama ≠ model training veya citation. bir verified hit -den GPTBot, ClaudeBot, veya PerplexityBot proves şu istek happened — bir fetch at şu layer. o yapmaz prove sayfa idi kullanılan -e train bir model, retained anywhere downstream, veya cited in bir chat yanıt. şunlar dır separate, unobserved outcomes; yapmayın stretch bir verified log line further -den o goes.

2026 wrinkle: AI bots dır tümü üzerinde sizin logs now

cast of characters in bir modern log file sahiptir changed. In benim analysis of Cloudflare Radar data (Meet yeni Web Crawlers), arama-motor bots hâlâ tarama en çok — ama AI bots dır firmly in ikinci place ve on izle -e overtake them bençinde bir couple of years. GPTBot, ClaudeBot, PerplexityBot, ve friends now göster up heavily. -dığınızda segment sizin verified hits tarafından kullanıcı agent, yapmayın olmak surprised -e bul AI crawlers rivaling arama motorlar bençin share of istekler. (Screaming Frog’s Log File Analyser sahiptir eklendi bir dedicated AI-bot tutorial bençin tam olarak bu.)

nasıl bu fits ile rest of tarama

Logs dır diagnostic layer altında whole tarama küme. onlar’re nasıl siz aslında measure tarama bütçesi spend şu motorlar describe in abstract (Gary Illyes defines o olarak ” number of URLs Googlebot -ebilir ve dır willing veya dır instructed -e tarama”). onlar’re ayrıca nasıl siz catch spider traps red-handed — bir infinite URL space -den bir calendar veya facet shows up olarak bir flood of near-identical istekler — ve nasıl siz yapğrula whether sizin tarama frequency çalışır (accurate lastmod, benç bağlantılar -e önemli sayfalar) aslında changed bot behavior. ve remember onlar complement, değil replace, GSC tarama istatistikler: tarama istatistikler dır sampled on-ramp; logs dır unsampled, multi-bot, per-URL detail.

Add an expert note

Pin an expert quote

New person? Create their unclaimed profile at /admin/experts/ → Pin a quote first.