Türkçe çeviri: Log File Analysis
nasıl -e okuyun sizin server's raw access logs -e see tam olarak ne Googlebot, Bingbot, ve AI crawlers aslında fetched — verifying gerçek bots, bulma tarama waste ve orphans, ve neden logs dır ground truth şu tarama araçlar ve arama Console yalnızca approximate.
Diller
Bu sayfada 1 kanıt sinyali
- Bağlantılı kaynak verileriyaygın-crawlers.json
Log file analysis dır okuma sizin server's raw access logs — unsampled, ground-truth record of her istek server aldı — -e see tam olarak hangi URLs Googlebot, Bingbot, ve AI crawlers aslında fetched, nasıl çoğu zaman, ve ile ne status code. non-negotiable ilk adım dır verifying bots dır gerçek (reverse + forward DNS, veya Google's published IP ranges), çünkü kullanıcı agents dır spoofed tümü time. o hâlde siz bak bençin tarama waste, en çok/least-tarandı URLs, status codes tarafından frequency, orphan sayfalar, ve mobile-vs-desktop split. Logs complement GSC tarama istatistikler; onlar yapmayın replace o. en çok küçük siteler yapmayın ihtiyaç duy bu — o's bir büyük-site, ecommerce, ve migration araç.
Evidence for this claim Web-server access logs record HTTP requests and commonly include request, response-status, user-agent, and timing fields depending on configuration. Scope: Apache HTTP Server access-log behavior; other servers vary by configuration. Confidence: high · Verified: Apache HTTP Server: Log Files Evidence for this claim User-agent text alone does not authenticate Googlebot; Google recommends DNS verification or matching published IP ranges. Scope: Google crawler verification, applicable when classifying log traffic. Confidence: high · Verified: Google Search Central: Verify Googlebot Evidence for this claim Cloudflare Radar compares worldwide Cloudflare-observed bot and human HTTP requests to HTML content during the 28 days ending 2026-07-30. Scope: A dated Cloudflare Radar context chart; the site's own verified logs remain the source of truth for site-specific traffic. Confidence: high · Verified: Cloudflare Radar: Bot versus human HTML trafficTL;DR — sizin web server keeps bir log of her istek o alır, dahil visits -den arama motoru bots like Googlebot. Log file analysis anlamına gelir okuma şu log -e see tam olarak hangi of sizin sayfalar bots aslında fetched, nasıl çoğu zaman, ve whether onlar hit errors. o’s yalnızca place şu shows ne gerçekten happened — ama en çok küçük siteler yapmayın ihtiyaç duy o.
The four-week chart compares automated and human requests to HTML content. Bot share is higher in the captured worldwide Cloudflare traffic period.
ne bir log file dır
her time anyone — bir kişben, Googlebot, veya bazı random bot — istekler bir sayfa -den sizin site, sizin server writes bir line hakkında o in bir file. şu file dır sizin access log. her line records kabaca aynı things:
- kim asked (bir IP adres ve bir “user agent” name like
Googlebot), - ne onlar asked bençin ( URL),
- ne zaman (bir timestamp), ve
- ne onlar aldı back (bir HTTP status code —
200bençin OK,404bençin değil found, ve bu nedenle on).
Stack up weeks of şunlar lines ve siz sahip bir complete, honest record of nasıl arama motorlar tarama sizin site. değil bir estimate. gerçek istekler.
neden bother okuma o
diğer SEO araçlar either guess at nasıl bots tarama siz (bir tarama araç pretends -e olmak bir arama motoru ve walks sizin site) veya summarize o (Google arama Console shows siz bir sampled, rounded-off view). sizin logs göster gerçek thing, istek tarafından istek. şu lets siz yanıt questions like:
- hangi sayfalar yapar Googlebot aslında visit — ve hangi yapar o ignore?
- dır bot wasting time on junk URLs yerine sizin önemli sayfalar?
- dır bots hitting broken sayfalar (
404) veya server errors (5xx)? - dır orada sayfalar bots sahip never ulaşılan at tümü?
bir kural -ebilirsiniz’t skip
Anyone -ebilir pretend -e olmak Googlebot. kullanıcı-agent name in bir log line dır sadece
text — bir scraper -ebilir put Googlebot in orada -e sneak past sizin defenses. bu nedenle
önce siz trust bir tek “Googlebot” line, siz sahip -e verify o’s gerçekten
Google (orada’s bir simple DNS kontrol et bençin bu in Advanced ve Scripts tabs).
Skip verification ve siz’ll draw conclusions -den fake trafik.
yap siz hatta ihtiyaç duy bu?
Honestly? Probably değil, -erseniz çalıştır bir küçük site. Log file analysis pays off bençin büyük siteler — tens of thousands of URLs, ecommerce ile lots of filtered sayfalar, siteler going aracılığıyla bir migration, veya siteler nerede Google şunu söylüyor o “discovered” sayfalar ama never dizine eklenmiş them. eğer sizin site sahiptir bir few hundred sayfalar ve onlar tümü al tarandı fine, sizin time dır daha iyi spent elsewhere. (aynı logic olarak tarama budget — en çok siteler yapmayın ihtiyaç duy -e worry hakkında o.)
iste gerçek workflow — verifying bots uygun biçimde, bulma tarama waste, spotting orphan sayfalar? Switch -e Advanced tab.
Evidence for this claim Web-server access logs record HTTP requests and commonly include request, response-status, user-agent, and timing fields depending on configuration. Scope: Apache HTTP Server access-log behavior; other servers vary by configuration. Confidence: high · Verified: Apache HTTP Server: Log Files Evidence for this claim User-agent text alone does not authenticate Googlebot; Google recommends DNS verification or matching published IP ranges. Scope: Google crawler verification, applicable when classifying log traffic. Confidence: high · Verified: Google Search Central: Verify GooglebotTL;DR — Logs dır unsampled ground truth bençin tarama: her istek, her bot, her status code. non-negotiable ilk adım dır verifying Googlebot/Bingbot via reverse + forward DNS (veya Google’s published IP-range JSON) — kullanıcı agents dır spoofed constantly, ve siz yap tümü tarama math on verified ayarla yalnızca. o hâlde okuyun logs bençin en çok/least-tarandı URLs ve sections, tarama frequency üzerinde time, status codes prioritized tarafından frequency, tarama waste (params, facets, site bençben arama, infinite sayfalama), orphan ve uncrawled sayfalar (cross-referenced karşı bir tarama), ve mobile-vs-desktop Googlebot split. Bing publishes no resmî IP ranges, bu nedenle DNS -e
*.search.msn.comdır yöntem orada. Logs complement GSC tarama istatistikler — onlar yapmayın replace o. ve in 2026, AI bots dır now bir huge share of ne shows up.
neden logs dır ground truth
vardır three ways -e “see” nasıl arama motorları tarama siz, ve onlar değildir equal:
- bir tarama araç (Screaming Frog SEO Spider, Ahrefs site denetimi) simulates bir tarama. o söyler siz ne bir bot -ebilirdi bul, değil ne Google yaptı fetch.
- GSC tarama istatistikler summarizes gerçek thing, ama o’s sampled, aggregated, ve capped (kabaca 1 000 rows, ~90 days, no per-URL export).
- Server logs record gerçek thing — her istek, bençin her bot, ile exact URL, timestamp, ve status code.
Ahrefs log-file rehber ben reviewed puts o plainly: server logs dır ” en çok trustworthy kaynak of information -e understand URLs şu arama motorları sahip tarandı.” şu’s whole neden bu technique vardır. ne zaman ben iste -e know ne Googlebot aslında yaptı — değil ne o -ebilir yap, değil bir rounded özet — ben go -e logs.
bir typical log line carries IP adres, kullanıcı agent, URL path, timestamp, istek yöntem (al/POST), ve HTTP status code. Everything below dır sadece slicing şunlar fields intelligently.
-dığınızda aslında ihtiyaç duy o (ve -dığınızda yapmayın)
olmak honest ile yourself burada. Log file analysis dır bir büyük-site araç. o earns onun
koru on siteler ile tens of thousands of URLs, ecommerce ve fasetli gezinme,
siteler mid-migration, ve siteler stuck in Discovered – currently not indexed. olarak ben
wrote in benim tarama bütçesi rehber, “en çok siteler yapmayın ihtiyaç duy -e worry hakkında tarama
budget, ama vardır few durumlar nerede siz -ebilir iste -e take bir bak.” Daniel
Waisberg of Google sahiptir yapılmış bir benzer benşaret et hakkında tarama istatistikler — per arama motoru
Journal’s coverage, rapor değildir much of bir concern bençin siteler altında ~1 000
sayfalar.
eğer sizin few-hundred-sayfa site dır olma tarandı fine, skip bu ve go düzelt something ile daha leverage.
nasıl -e al sizin logs ( hardest part dır access)
Logs live wherever istek aslında terminated:
- Apache & Nginx → Apache “combined” log format ( en çok yaygın).
- Microsoft IIS → W3C format.
- AWS ELB/ALB → ELB format.
- CDNs (Cloudflare, Fastly, Akamai) → onların kendi log exports. bu önem taşır: at bir CDN-fronted site, bir origin-yalnızca log misses edge-cached hits, bu nedenle pull logs at layer bot aslında ulaşılan.
Aim bençin 30 days minimum, 90 ideal, -e capture tarama-frequency variation. ve plan bençin friction — getting access -e server logs dır çoğu zaman genuinely hard part (DevOps gatekeeping). hatta Googlers, in bir migration episode of arama Off Record, flagged nasıl hard log files -ebilir olmak -e obtain uygulamada. Budget time bençin istek.
Logs değildir sadece bot trafik — onlar capture her istek, dahil gerçek visitors, ve -ebilir carry sorgu-string values, session identifiers, veya diğer sensitive data alongside URL path. OWASP’s logging rehberlik dır blunt hakkında bu: authentication credentials, access tokens, ve personally identifiable information generally shouldn’t land yapğrudan in bir log; onlar -meli olmak removed, masked, veya hashed ilk. oluştur şu -e sizin access controls ve export süreç önce siz hand bir log file -e anyone bençin analysis, değil sonra.
adım 1 — Verify bots dır gerçek (yap bu önce anything else)
bu adım en çok guides wave at in bir line. yapmayın. çok söyleıda bots pretend -e olmak Googlebot -e al past firewalls (Ahrefs). kullanıcı agent dır unauthenticated text; ele al her “Googlebot” line olarak bir claim -e olmak proven.
Googlebot — iki geçerli yöntem:
- Reverse + forward DNS ( bidirectional kontrol et). Google’ın kendi adımlar: çalıştır bir
reverse DNS lookup on IP -den sizin logs ile
hostcommand; verify domain dırgooglebot.com,google.com, veyagoogleusercontent.com; o hâlde çalıştır bir forward DNS lookup on şu hostname ve verify o resolves back -e özgün IP. forward adım dır ne yapar bu trustworthy — bir spoofer -ebilir benşaret et reverse DNS at bir*.googlebot.comname, ama yalnızca round-trip back -e aynı IP proves o. (Commands bençin macOS/Linux ve Windows dır in Scripts tab.) - Match karşı Google’s published IP ranges. Google publishes JSON files of
onun crawler IPs in CIDR format —
common-crawlers.jsonbençin Googlebot ve friends, plusspecial-crawlers.json, kullanıcı-triggered-fetcher files, ve bir tümü-Googlegoog.json. olarak ben noted in benim Googlebot rehber, Google “provided bir liste of kamuya birçık IPs -ebilirsiniz kullan -e verify istekler dır -den Google… -ebilirsiniz compare bu -e data in sizin server logs.”
Bingbot — DNS yalnızca. bu key contrast: Bing yapmaz officially
publish IP ranges. Bing’s kendi wording dır şu “…like diğer arama motorları,
Bing yapmaz publish bir liste of IP adresler veya ranges -den hangi biz tarama
Internet,” because ” IP adresler veya ranges biz kullan -ebilir change herhangi bir time.” bu nedenle
bençin Bingbot siz yap reverse + forward DNS -e bir hostname ending in
*.search.msn.com (e.g. msnbot-157-55-33-18.search.msn.com), veya kullan
Verify Bingbot araç. (Microsoft
sahiptir since released bir bingbot IP JSON, ama onun resmî verification rehberlik
hâlâ centers on DNS precisely çünkü IPs change.)
o hâlde throw out fakes. yap tümü of sizin tarama math on verified ayarla yalnızca. Unverified “Googlebot” dır almost her zaman bir scraper veya bir spoofed bot ve belongs in bir security review, değil sizin tarama-waste analysis.
adım 2 — ne -e bak bençin
Once siz’re working ile verified hits, burada’s okuyun:
- en çok & least tarandı URLs ve sections. Rank istekler tarafından URL ve tarafından directory. bu nerede sizin tarama bütçesi dır aslında going — ve o’s genellikle surprising.
- tarama frequency üzerinde time. Trend tarar tarafından URL/section -e catch drops (bir migration broke something) veya spikes (bir yeni section, veya bir spider trap spinning up infinite URLs).
- Status codes bots hit, prioritized tarafından frequency. Quantify
200vs.301/302(ve chains),404, ve5xx. bir404hit 5 000×/week dır bir farklı sorun -den bir404hit once — düzelt tarafından tarama frequency, değil tarafından mere existence. - tarama waste. Faceted nav, URL parametreleri, site bençben arama sonuçlar, ve infinite calendars/sayfalama -ebilir eat bir büyük share of tarama bütçesi on bad offenders. Logs göster tam olarak hangi junk patterns bots dır burning time on.
- Orphan & uncrawled sayfalar. bu gerektirir her ikisi datasets. Cross-reference logs karşı bir site tarama: URLs in logs ama değil in tarama = orphans, eski yönlendirmeler, veya externally-linked sayfalar; URLs in tarama ama değil in logs = sayfalar Google sahiptir never fetched.
- Mobile vs. desktop Googlebot. Split tarafından kullanıcı agent. Post mobile-ilk, o -meli olmak majority Googlebot Smartphone — bir desktop-heavy split dır worth bir bak.
- Response time & tarama health. Rising average response time correlates ile reduced tarama. Per SEJ’s yaz-up of Waisberg’s rehberlik: “Watch out bençin bir consistent increase in average response time. Google şunu söylüyor o -ebilir değil affect tarama rate immediately, ama o’s bir good indicator şu sizin servers -ebilir değil olmak handling tümü load.”
ne logs yapmayın söyle siz
koru bunlar straight veya siz’ll üzerinde-okuyun data:
- tarama ≠ dizin. bir URL Googlebot fetches daily -ebilir stay unindexed indefinitely. Logs prove fetching, değil dizin status — pair them ile GSC’s sayfa dizine ekleme / URL Inspection -e öğren dizin side.
- tarama ≠ rank, ve daha tarama yapmaz yardım et. olarak ben’ve said repeatedly, “The rate of crawling isn’t going to impact your rankings.” yapmayın chase tarama volume olarak eğer o idi bir sıralama lever.
noindexyapmaz reduce tarama.noindexcontrols indexation, değil tarama — -e aslında durdur tarama siz kullan robots.txt veya bir status code.- tarama ≠ model training veya citation. bir verified hit -den GPTBot, ClaudeBot, veya PerplexityBot proves şu istek happened — bir fetch at şu layer. o yapmaz prove sayfa idi kullanılan -e train bir model, retained anywhere downstream, veya cited in bir chat yanıt. şunlar dır separate, unobserved outcomes; yapmayın stretch bir verified log line further -den o goes.
2026 wrinkle: AI bots dır tümü üzerinde sizin logs now
cast of characters in bir modern log file sahiptir changed. In benim analysis of Cloudflare Radar data (Meet yeni Web Crawlers), arama-motor bots hâlâ tarama en çok — ama AI bots dır firmly in ikinci place ve on izle -e overtake them bençinde bir couple of years. GPTBot, ClaudeBot, PerplexityBot, ve friends now göster up heavily. -dığınızda segment sizin verified hits tarafından kullanıcı agent, yapmayın olmak surprised -e bul AI crawlers rivaling arama motorlar bençin share of istekler. (Screaming Frog’s Log File Analyser sahiptir eklendi bir dedicated AI-bot tutorial bençin tam olarak bu.)
nasıl bu fits ile rest of tarama
Logs dır diagnostic layer altında whole tarama küme. onlar’re nasıl
siz aslında measure tarama bütçesi spend şu motorlar describe in
abstract (Gary Illyes defines o olarak ” number of URLs Googlebot -ebilir ve dır
willing veya dır instructed -e tarama”). onlar’re ayrıca nasıl siz catch spider traps
red-handed — bir infinite URL space -den bir calendar veya facet shows up olarak bir flood of
near-identical istekler — ve nasıl siz yapğrula whether sizin tarama frequency
çalışır (accurate lastmod, benç bağlantılar -e önemli sayfalar) aslında changed bot
behavior. ve remember onlar complement, değil replace, GSC tarama istatistikler: tarama
istatistikler dır sampled on-ramp; logs dır unsampled, multi-bot, per-URL detail.
AI özet
bir condensed take on Advanced sürüm:
- Logs = ground truth. tarama araçlar simulate, GSC samples; server access logs record her istek, her bot, ile URL, timestamp, ve status code.
- Verify önce siz analyze. kullanıcı agents dır spoofed constantly. yapğrula
Googlebot via reverse + forward DNS veya Google’s published IP-range JSON; yapğrula
Bingbot via reverse DNS -e
*.search.msn.com(Bing publishes no resmî IP ranges). yap tümü tarama math on verified ayarla yalnızca. - ne -e okuyun: en çok/least-tarandı URLs ve sections; tarama frequency üzerinde time; status codes prioritized tarafından frequency (bir 404 hit 5 000×/week ≠ once); tarama waste (params, facets, site bençben arama, infinite sayfalama); orphan ve uncrawled sayfalar (cross-reference bir tarama); mobile-vs-desktop Googlebot split; rising response time olarak bir tarama-health warning.
- yapmayın üzerinde-okuyun o: tarama ≠ dizin ≠ rank, daha tarama yapmaz yardım et
sıralamalar,
noindexyapmaz reduce tarama, ve bir verified AI-bot hit proves bir fetch — değil şu sayfa idi kullanılan bençin model training veya cited in bir yanıt. - Scope: bir büyük-site / ecommerce / migration araç. en çok küçük siteler yapmayın ihtiyaç duy o.
- 2026: AI bots (GPTBot, ClaudeBot, PerplexityBot) dır now bir major ve growing share of log trafik.
- Complements GSC tarama istatistikler — kullan her ikisi.
resmî dokümantasyon
birincil-kaynak dokümantasyon bençin verifying crawlers ve okuma tarama data.
- Verify istekler -den Google Crawlers ve Fetchers — two resmî methods: manual reverse + forward DNS, ve matching karşı published IP ranges.
- nasıl -e verify Googlebot (arama Central Blog) — older companion post, hâlâ cited.
- yaygın-crawlers.json — Googlebot ve yaygın crawler IPs in CIDR format (note: “the IP addresses in the JSON files are represented in CIDR format”). Companions: special-crawlers.json, kullanıcı-triggered-fetchers.json, ve tümü-Google goog.json.
- Overview of Google crawlers ve fetchers — her Google kullanıcı agent siz’ll see in logs.
- tarama istatistikler rapor (yardım et) ve launch blog — sampled, resmî view şu logs complement.
Bing / Microsoft
- nasıl -e Verify Bingbot (Webmaster yardım et) — resmî verification sayfa.
- nasıl -e Verify şu Bingbot dır Bingbot (blog) — canonical reverse + forward DNS yöntem, ve statement şu Bing yapmaz publish IP ranges.
- Verify Bingbot araç — paste bir IP -e kontrol et o.
Quotes -den kaynak
On—record statements. nerede supported, her bağlantı dır bir deep bağlantı şu jumps -e quoted passage on kaynak sayfa.
Google — tarayıcıları doğrulama
- “Run a reverse DNS lookup on the accessing IP address from your logs, using the
hostcommand.” — Google arama Central docs. Jump -e quote - “Verify that the domain name is either
googlebot.com,google.com, orgoogleusercontent.com.” Jump -e quote - “Run a forward DNS lookup on the domain name retrieved in step 1 using the
hostcommand on the retrieved domain name.” Jump -e quote - “Verify that it’s the same as the original accessing IP address from your logs.” Jump -e quote
Bing — verification ve no published IP ranges
- “Perform a reverse DNS lookup using the IP address from the logs to verify that it resolves to a name that end with search.msn.com.” — Bing Webmaster Blog. Jump -e quote
- “…like other search engines, Bing does not publish a list of IP addresses or ranges from which we crawl the Internet.” — ve neden: “the IP addresses or ranges we use can change any time, so responding to requests differently based on a hardcoded list is not a recommended approach.” Jump -e quote
Patrick Stox — on ne logs dır bençin (-den benim çalışır at Ahrefs)
- Server logs dır “the most trustworthy source of information to understand the URLs that search engines have crawled.” Jump -e quote
- “Many bots pretend to be Googlebot to get past firewalls.” Jump -e quote
- “If you want to see hits from all bots and users, you’ll need access to your log files.” Jump -e quote
- “The rate of crawling isn’t going to impact your rankings.” Jump -e quote
Gary Illyes, Google — tarama bütçesi (ne logs izin ver siz measure)
- tarama bütçesi dır “the number of URLs Googlebot can and is willing or is instructed to crawl.” okuyun coverage
Daniel Waisberg, Google — response time olarak bir tarama-health sinyal (SEJ’s framing of his tarama istatistikler rehberlik)
- “Watch out for a consistent increase in average response time. Google says it might not affect crawl rate immediately, but it’s a good indicator that your servers might not be handling all the load.” Jump -e quote
Verify bir bot dır gerçekten Googlebot (reverse + forward DNS)
kullanıcı agent in bir log line dır sadece text — scrapers spoof Googlebot -e al past
firewalls. yalnızca trustworthy kontrol et dır bidirectional DNS: reverse-lookup IP,
yapğrula hostname dır bir Google domain, o hâlde forward-lookup şu hostname ve
yapğrula o resolves back -e aynı IP.
macOS / Linux (kullanır host)
# 1) Reverse DNS the IP from your logs — it must end in googlebot.com,
# google.com, or googleusercontent.com
host 66.249.66.1
# → 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com
# 2) Forward DNS that hostname back — it must resolve to the same IP
host crawl-66-249-66-1.googlebot.com
# → crawl-66-249-66-1.googlebot.com has address 66.249.66.1Windows (kullanır nslookup)
:: 1) Reverse DNS the IP — confirm it ends in a Google domain
nslookup 66.249.66.1
:: 2) Forward DNS the returned hostname — confirm it matches the original IP
nslookup crawl-66-249-66-1.googlebot.comeğer reverse lookup yapmaz land on bir Google domain, veya forward lookup
yapmaz döndür özgün IP, o dır değil Googlebot. (bençin Bingbot, çalıştır
exact aynı two adımlar ama expect bir hostname ending in search.msn.com.) -ebilirsiniz
ayrıca skip DNS ve match IP karşı Google’s published ranges
(yaygın-crawlers.json,
in CIDR format).
Extract ve count crawler hits -den bir raw log
Quick bir-liners bençin bir combined-format Apache/Nginx access log. (bunlar seç on kullanıcı-agent string — remember -e verify IPs önce trusting counts.)
# Pull only the lines claiming to be Googlebot
grep -i "googlebot" access.log > googlebot-hits.log
# Count Googlebot hits per URL, most-crawled first
# (combined format: $7 is the request path)
grep -i "googlebot" access.log \
| awk '{print $7}' \
| sort | uniq -c | sort -rn | head -50
# Count Googlebot hits per status code (combined format: $9 is the status)
grep -i "googlebot" access.log \
| awk '{print $9}' \
| sort | uniq -c | sort -rn
# List the URLs Googlebot hit that returned a 404, by frequency
grep -i "googlebot" access.log \
| awk '$9 == 404 {print $7}' \
| sort | uniq -c | sort -rn
# Get the unique IPs claiming Googlebot — the list you then verify by DNS
grep -i "googlebot" access.log | awk '{print $1}' | sort -uAdjust $7/$9 field positions eğer sizin log format differs (IIS/W3C ve ELB
sıra fields differently).
Günlük dosyası analiz iş akışı
- al logs — server access logs, veya CDN/load-balancer exports. At bir CDN-fronted site, pull edge logs de (origin misses cache hits).
- Grab enough window — 30 days minimum, 90 ideal.
- Verify bots ilk — reverse + forward DNS, veya Google’s published
IP ranges. Bingbot → DNS -e
*.search.msn.com. - Drop fakes — yap tümü tarama math on verified ayarla; route spoofed “Googlebot” -e bir security review.
- Rank en çok/least-tarandı URLs ve sections — see nerede budget goes.
- Trend tarama frequency üzerinde time — catch drops (migrations) ve spikes (yeni sections, spider traps).
- Tally status codes —
200/301-302chains /404/5xx, ve prioritize düzeltmeler tarafından tarama frequency, değil mere existence. - Hunt tarama waste — parameters, faceted nav, site bençben arama, infinite sayfalama/calendars.
- bul orphans & uncrawled sayfalar — cross-reference logs karşı bir site tarama (in logs değil tarama = orphan; in tarama değil logs = never fetched).
- kontrol et mobile vs. desktop Googlebot split — -meli olmak Smartphone-majority.
- Watch average response time — bir rising trend -ebilir throttle tarama.
- Segment AI bots — GPTBot, ClaudeBot, PerplexityBot, etc. dır now bir gerçek share of trafik.
- yapğrula ile GSC — pair tarama findings ile tarama istatistikler ve URL Inspection (tarama ≠ dizin).
ne -e oluştur in sizin analysis sheet
Whether siz kullan Screaming Frog’s Log File Analyser, BigQuery, veya Ahrefs log-file template, siz’re building aynı handful of views. sonra verifying bots, parse her line -e columns ve pivot.
Columns -e parse out of her log line
| Column | -den log line | neden önemli olduğu |
|---|---|---|
| IP adres | $1 (combined format) | thing siz verify tarafından DNS / IP range |
| Verified? | derived | Filter her pivot -e verified yalnızca |
| Bot / kullanıcı agent | UA string | Segment Googlebot Smartphone vs. Desktop vs. AI bots |
| URL path | $7 | grup tarar tarafından URL ve tarafından directory |
| Section / directory | derived -den path | Roll up tarama spend tarafından site area |
| Timestamp | [date] field | Trend tarama frequency üzerinde time |
| yöntem | al/POST | Spot odd istek patterns |
| Status code | $9 | 200 / 3xx / 404 / 5xx breakdown |
Pivots -e oluştur
- tarar per URL (ve per directory), descending — en çok & least tarandı.
- tarar per status code, o hâlde status code × URL (bu nedenle bir high-frequency
404jumps out). - tarar per day/week, segmented tarafından section — frequency trend.
- tarar per bot/kullanıcı agent — mobile/desktop split, plus AI-bot share.
- bir logs vs. tarama join (VLOOKUP/merge karşı bir Screaming Frog veya Ahrefs tarama export) -e surface orphans ve never-fetched sayfalar.
Ahrefs rehber ships bir downloadable template şu sets bunlar up; log-file-analysis rehber bağlantılar o.
nasıl -e okuyun bir log file: ne -e bak bençin
bir repeatable lens bençin herhangi bir log analysis. Verify ilk; o hâlde çalıştır bunlar geçer.
1. Verify, o hâlde analyze. data dır yalnızca olarak good olarak bot identity behind o. Reverse + forward DNS (veya IP ranges); analyze verified ayarla yalnızca. Spoofed bots → security, değil SEO.
2. nerede dır budget going? (en çok vs. least tarandı.) Rank tarafından URL ve tarafından section. goal dır -e bul attention spent on yanlış places — ve önemli sayfalar getting de little.
3. ne dır bots hitting? (status codes, tarafından frequency.)
200 dır healthy; 3xx chains, 404, ve 5xx dır leaks. Triage tarafından nasıl çoğu zaman
bot hits her, değil tarafından whether error merely vardır.
4. ne’s olma wasted? (tarama waste.) Parameters, faceted nav, site bençben arama, infinite sayfalama/calendars — classic budget sinks. Logs name exact offending patterns.
5. ne’s missing? (orphans & uncrawled.) Cross-reference logs ile bir tarama. In logs ama değil tarama = orphan/eski yönlendirme/external bağlantı. In tarama ama değil logs = Google never fetched o.
6. kim’s tarama? (bot segmentation.) Mobile vs. desktop Googlebot (-meli olmak Smartphone-majority), plus AI-bot share şu now rivals arama motorları.
7. dır server healthy? (response time.) Rising average response time dır bir early warning şu tarama -ebilir al throttled.
araçlar bençin log file analysis
- Screaming Frog SEO Log File Analyser — workhorse desktop araç. o auto-verifies arama bots ve flags spoofed IPs, ve sahiptir kullanıcı-Agent + Verification-Status filters bençin granular bot analysis. Import bir SEO Spider tarama ve kullan “Not In URL Data” filter -e bul orphans — “URLs hangi idi discovered in sizin logs, ama değildir present in tarama data imported.” o ayrıca now sahiptir bir dedicated AI-bot izleme tutorial. araç
- BigQuery (ve Splunk / ELK Stack / Logflare / logz.io) — bençin raw, büyük-scale storage ve querying ne zaman log volume dır de big bençin bir desktop araç.
- Ahrefs — cross-reference logs karşı bir Ahrefs site denetimi tarama -e bul orphans ve yapğrula ne bots ulaşılan. (benim home araç.)
- Semrush Log File Analyzer, OnCrawl, Botify, JetOctopus — diğer tarama+log platforms, latter three aimed at enterprise.
- GSC tarama istatistikler — resmî, free, sampled on-ramp. Complementary -e raw logs, değil bir replacement: o’s aggregated, capped, ve sahiptir no per-URL export.
- Verify Bingbot araç — paste bir IP -e yapğrula o’s gerçekten Bingbot: bing.com/toolbox/verify-bingbot.
Mistakes -e kaçın
- Trusting “Googlebot” string olmadan verifying o. neden o’s yanlış:
kullanıcı agent dır unauthenticated text — anyone -ebilir put
Googlebotin bir istek header -e slip past defenses veya pollute sizin analysis ile fake trafik. yap instead: çalıştır reverse + forward DNS (veya match karşı Google’s published IP ranges) önce bir tek line counts toward sizin tarama math. - Treating GSC tarama istatistikler olarak full picture. neden o’s yanlış: o’s sampled, rounded, capped at kabaca 1 000 rows ve ~90 days, ile no per-URL export — o summarizes, o yapmaz record. yap instead: kullan logs olarak ground truth ve tarama istatistikler olarak bir complementary, faster on-ramp.
- Pulling origin-yalnızca logs on bir CDN-fronted site. neden o’s yanlış: bir origin log misses her istek CDN sunulan -den edge cache, bu nedenle siz’re analyzing bir incomplete picture of ne bots aslında aldı. yap instead: pull logs at layer bot aslında ulaşılan — CDN/edge export, değil sadece origin server.
- düzeltme 404s ve 5xx errors tarafından existence, değil frequency. neden o’s yanlış:
bir
404hit once bir month ve bir404hit 5 000 times bir week değildir aynı sorun, ama treating her error line olarak equally urgent wastes düzelt effort. yap instead: rank status-code sorunlar tarafından nasıl çoğu zaman bots aslında hit them. - Chasing tarama volume olarak eğer o idi bir sıralama lever. neden o’s yanlış: daha tarama yapmaz move sıralamalar — o’s bir diagnostic sinyal, değil bir growth metric. yap instead: kullan tarama frequency -e catch problems (drops, spider traps), değil olarak bir KPI -e maximize.
- kullanarak
noindex-e try -e durdur tarama. neden o’s yanlış:noindexcontrols indexation, değil tarama — Google hâlâ sahiptir -e fetch sayfa -e see tag. yap instead: block tarama itself ile robots.txt veya bir status code eğer şu’s gerçek goal. - Calling bir URL “orphaned” -den logs alone. neden o’s yanlış: bir URL şu shows up in logs ama değil in bir fresh tarama -ebilirdi olmak bir eski yönlendirme, bir external bağlantı, veya bir genuine orphan — logs alone -ebilir’t söyle siz hangi. yap instead: cross-reference logs karşı bir gerçek site tarama önce drawing conclusions either way.
Standing KPIs bençin log file analysis
Doğrulanmış bot payı
- Metric: Verified Googlebot/Bingbot hits ÷ tümü hits claiming -e olmak Googlebot/Bingbot tarafından kullanıcı agent.
- ne o söyler siz: nasıl much of sizin “bot traffic” dır aslında spoofed scrapers yerine gerçek arama motorları.
- nasıl -e pull o: çalıştır her claimed-bot IP aracılığıyla reverse + forward DNS (veya IP-range JSON) ve count geç rate.
- Benchmark / realistic range: No universal number — depends on nasıl aggressively sizin site dır scraped. bir low veya dropping verified share dır sinyal -e act on, değil bir düzeltilmiş threshold.
- Cadence: her time siz pull bir yeni log window.
tarama waste share
- Metric: % of verified bot istekler hitting parameters, facets, internal arama, veya infinite sayfalama/calendars.
- ne o söyler siz: nasıl much of sizin tarama bütçesi dır going -e junk URLs yerine sayfalar şu önem taşır.
- nasıl -e pull o: Segment verified hits tarafından URL pattern (sorgu strings, known facet/arama paths).
- Benchmark / realistic range: No universal figure — genuinely situational, depending on sizin site’s URL structure ve faceting. Establish sizin kendi baseline on ilk pull, o hâlde izle trend.
- Cadence: 30–90 day log window; re-kontrol et sonra herhangi bir cleanup (robots.txt kurallar, parametre yönetimi, sayfalama düzeltmeler).
Status-code mix, weighted tarafından frequency
- Metric:
200/3xx/404/5xxolarak bir share of verified bot istekler. - ne o söyler siz: nerede bots dır burning fetches on errors yerine live bençerik, ve whether şu’s getting worse.
- nasıl -e pull o: Tally status-code field -den verified log lines.
- Benchmark / realistic range: No universal target — depends on site age
ve yönlendirme history. Watch trend, değil bir tek snapshot; bir rising
404/5xxshare dır actionable sinyal. - Cadence: 30–90 days, veya immediately sonra bir migration.
Mobil ve masaüstü Googlebot dağılımı
- Metric: Share of verified Googlebot hits -den Smartphone kullanıcı agent vs. Desktop.
- ne o söyler siz: Whether Google dır aslında tarama siz mobile-ilk, olarak beklenen post mobile-ilk dizine ekleme.
- nasıl -e pull o: Segment verified hits tarafından Googlebot UA string (Smartphone vs. Desktop).
- Benchmark / realistic range: -meli olmak Smartphone-majority bençin en çok siteler; bir desktop-heavy split dır worth investigating, değil bir hard başarısızlık on onun kendi.
- Cadence: her log pull.
Orphan / uncrawled sayfa count
- Metric: URLs in logs ama değil in bir fresh tarama (orphans/eski bağlantılar), ve URLs in bir fresh tarama ama değil in logs (never fetched).
- ne o söyler siz: sayfalar bots -ebilir’t easily ulaş, ve sayfalar siz’re linking -e şu Google sahiptir never bothered -e fetch.
- nasıl -e pull o: Join verified-hit URL listesi karşı bir site tarama export (Screaming Frog, Ahrefs).
- Benchmark / realistic range: Entirely situational — depends on site size ve nasıl recently siz migrated veya restructured. izle count üzerinde time yerine comparing -e bir outside number.
- Cadence: Quarterly bençin büyük siteler, veya immediately sonra bir migration.
Ready—e-kopya prompts bençin log analysis
bunlar dır bençin interpreting log data siz’ve zaten pulled ve verified — değil bençin generating log data (never izin ver bir AI invent log lines veya istatistikler).
Summarize tarama waste -den bir URL sample
Here is a list of URL paths that verified Googlebot hits landed on, one per
line, from my server logs. Group them into patterns (query parameters,
faceted navigation, internal search, pagination/calendars, or "looks like a
real page"), and tell me which pattern has the most URLs. Don't invent URLs
that aren't in the list — only group what I've pasted.
[paste your URL list here]What to expect back: a small set of named clusters with counts, plus a flag on which cluster looks like the biggest crawl-waste offender — treat it as a starting point to verify against the actual paths, not a final answer.
Prioritize bir status-code breakdown
I have this table of HTTP status codes and how many times verified Googlebot
hit each one over the last 30 days. Rank them by which I should fix first,
weighting frequency over severity — a 404 hit 5,000 times matters more than a
500 hit twice. Explain the reasoning in one line per row.
status_code, hit_count
[paste your table here]What to expect back: the same rows re-ordered by fix priority with a one-line reason each — sanity-check the reasoning against your own site context before acting.
Draft bir log-access istek -e DevOps
Write a short, plain-English email to my DevOps/hosting team asking for
30-90 days of raw web server access logs (Apache/Nginx combined format, or
our CDN's edge logs if we're behind one) for [site name]. Explain in one
sentence why I need it (verifying real Googlebot/Bingbot crawl activity vs
GSC's sampled report) and ask what export format and delivery method works
for them.What to expect back: a short draft email you can edit with your actual site name and send — review it yourself before sending, this tool won’t send messages on your behalf.
Explain bir DNS verification sonuç
I ran a reverse DNS lookup on an IP from my server logs and then a forward
DNS lookup on the hostname it returned. Here's the raw output from the
`host` command. Tell me plainly whether this confirms the request came from
real Googlebot or Bingbot, and point to exactly which line proves or
disproves it.
[paste your host/nslookup output here]What to expect back: a plain-language verdict tied to the specific line in your output — treat it as a second opinion, not a replacement for knowing the actual rule (hostname ends in a Google/Bing domain, and the forward lookup returns the original IP).
kaynaklar worth sizin time
benim related yazma
- nasıl -e yap bir SEO Log File Analysis [Template Included] — Ahrefs cornerstone rehber ben reviewed; framing ve template bu article leans on.
- ne zaman -meli siz Worry hakkında tarama bütçesi? — ne zaman log analysis dır (ve değildir) worth sizin time.
- ne dır Googlebot & nasıl yapar o çalışır? — crawlers ve IP-verification background.
- Meet yeni Web Crawlers: AI Bots dır Closing in on arama motoru Bots — neden sizin logs bak farklı in 2026.
- Beginner’s rehber -e teknik SEO — nerede tarama ve logs fit in bigger picture.
benim speaking
- nasıl arama çalışır (SlideShare) — benim walkthrough of tarama: Googlebot olarak 1 000+ systems ile çok söyleıda specialized crawlers (Desktop, Mobile, Image, News, Video, Ads) sharing bir tarama-budget pool, ve istekler mostly originating out of Mountain View — bir handy log sanity-kontrol et alongside, never yerine, proper verification. (Standing disclaimer: “This is my understanding of systems… not going to be 100% complete or accurate.”)
-den others
- Screaming Frog Log File Analyser — kullanıcı rehber ve AI-bot tutorial.
- arama motoru Land — Log file analysis rehber (Kody Wirth) — practical walkthrough covering formats, araçlar, ve ne patterns -e bak bençin.
- arama motoru Journal — nasıl -e kullan Google’s tarama istatistikler rapor — Daniel Waisberg’s (Google) rehberlik on okuma tarama istatistikler, dahil response-time warning; en iyi companion -e log analysis.
- arama motoru Journal — dizine ekleme & tarama bütçesi (Illyes + Splitt) — on—record definition of tarama bütçesi ve nasıl quality drives tarama demand.
- arama motoru Roundtable — Bingbot IP adresler released — kapsar nuance şu Microsoft eventually published bir bingbot IP JSON hatta gerçben resmî rehberlik hâlâ centers on DNS.
- Conductor — Log File Analysis bençin SEO — accessible explainer of ne logs reveal ve nasıl -e act on them.
- r/TechSEO — community bençin tarama/dizin debugging.
Videos
- Google arama Central (YouTube) — Martin Splitt’s tarama/rendering explainers ve nasıl Google arama çalışır series; yararlı background on bots whose hits siz’re verifying in sizin logs. Channel
test et yourself
Five questions on verifying bots ve okuma ne sizin logs aslında göster.
Değişiklik günlüğü
8 Ağu 2026 tarihinde güncellendi.
Editoryal özet ve kaydedilen değişiklik ayrıntıları.Değişiklik ayrıntıları
-
Ayrıntılı değişiklik notları şu anda İngilizce olarak mevcut.
Tam karşılaştırma kullanılamıyor — bu sürüm için önceki anlık görüntü arşivlenmemiş.
30 Tem 2026 tarihinde güncellendi.
Editoryal özet ve kaydedilen değişiklik ayrıntıları.Değişiklik ayrıntıları
-
Ayrıntılı değişiklik notları şu anda İngilizce olarak mevcut.
Tam karşılaştırma kullanılamıyor — bu sürüm için önceki anlık görüntü arşivlenmemiş.
18 Tem 2026 tarihinde güncellendi.
Editoryal özet ve kaydedilen değişiklik ayrıntıları.Değişiklik ayrıntıları
-
Ayrıntılı değişiklik notları şu anda İngilizce olarak mevcut.
-
Ayrıntılı değişiklik notları şu anda İngilizce olarak mevcut.
-
Ayrıntılı değişiklik notları şu anda İngilizce olarak mevcut.
Tam karşılaştırma kullanılamıyor — bu sürüm için önceki anlık görüntü arşivlenmemiş.