Log File phân tích
Cách đọc máy chủ của bạn thô access logs để see chính xác điều gì Googlebot, Bingbot, và AI các crawler thực ra fetched — verifying real bots, finding crawl waste và orphans, và vì sao logs là đó ground truth đó crawl tools và Search Console chỉ approximate.
Ngôn ngữ
1 tín hiệu bằng chứng trên trang này
- Dữ liệu nguồn được liên kếtphổ biến-các crawler.json
Log file analysis là reading máy chủ của bạn thô access logs — đó unsampled, ground-truth record of mỗi yêu cầu đó máy chủ đã nhận — để see chính xác mà URLs Googlebot, Bingbot, và AI các crawler thực ra fetched, cách thường, và với điều gì mã trạng thái. Đó non-negotiable đầu tiên step là verifying đó bots là real (reverse + forward DNS, hoặc Google published IP ranges), vì người dùng agents là spoofed all đó time. Thì bạn tìm crawl waste, hầu hết/least-được crawl URLs, các mã trạng thái by frequency, orphan các trang, và đó mobile-so với-desktop split. Logs complement GSC Crawl Số liệu; they không replace điều này. Hầu hết nhỏ các trang không cần này — đây là một lớn-site, ecommerce, và migration tool.
Evidence for this claim Web-server access logs record HTTP requests and commonly include request, response-status, user-agent, and timing fields depending on configuration. Scope: Apache HTTP Server access-log behavior; other servers vary by configuration. Confidence: high · Verified: Apache HTTP Server: Log Files Evidence for this claim User-agent text alone does not authenticate Googlebot; Google recommends DNS verification or matching published IP ranges. Scope: Google crawler verification, applicable when classifying log traffic. Confidence: high · Verified: Google Search Central: Verify Googlebot Evidence for this claim Cloudflare Radar compares worldwide Cloudflare-observed bot and human HTTP requests to HTML content during the 28 days ending 2026-07-30. Scope: A dated Cloudflare Radar context chart; the site's own verified logs remain the source of truth for site-specific traffic. Confidence: high · Verified: Cloudflare Radar: Bot versus human HTML trafficTL;DR — Your web server keeps a log of every request it gets, including visits from search engine bots like Googlebot. Log file analysis means reading that log to see exactly which of your pages the bots actually fetched, how often, and whether they hit errors. It’s the only place that shows what really happened — but most small sites don’t need it.
The four-week chart compares automated and human requests to HTML content. Bot share is higher in the captured worldwide Cloudflare traffic period.
Điều gì một log file là
Mỗi khi anyone — một person, Googlebot, hoặc some random bot — các yêu cầu một trang từ trang web của bạn, máy chủ của bạn ghi một line về điều này trong một file. Đó file là của bạn access log. Mỗi line records khoảng đó giống nhau điều:
- ai asked (an IP address và một “người dùng agent” name như
Googlebot), - điều gì they asked cho (đó URL),
- khi (một timestamp), và
- điều gì they đã nhận lại (an HTTP mã trạng thái —
200cho OK,404cho không được tìm thấy, và so on).
Stack lên weeks of những lines và bạn có một hoàn tất, honest record of cách tìm kiếm engines crawl trang web của bạn. Không an estimate. Đó thực tế các yêu cầu.
Vì sao bother reading điều này
Other SEO tools either guess tại cách bots crawl bạn (một crawl tool pretends để là một công cụ tìm kiếm và walks trang web của bạn) hoặc summarize điều này (Google Search Console cho thấy bạn một sampled, rounded-off view). Của bạn logs cho thấy đó real điều, yêu cầu by yêu cầu. Đó cho phép bạn câu trả lời các câu hỏi như:
- Mà các trang làm Googlebot thực ra visit — và mà làm điều này bỏ qua?
- Là đó bot wasting time on junk URLs thay vì của bạn quan trọng các trang?
- Là bots hitting hỏng các trang (
404) hoặc máy chủ các lỗi (5xx)? - Là ở đó các trang bots có không bao giờ reached tại all?
Đó một rule bạn không thể skip
Anyone có thể pretend để là Googlebot. Người dùng-agent name trong một log line là chỉ
text — một scraper có thể put Googlebot trong ở đó để sneak past của bạn defenses. So
trước khi bạn trust một single “Googlebot” line, bạn có để verify đây là thực sự
Google (có một đơn giản DNS kiểm tra cho này trong đó Advanced và Scripts tabs).
Skip verification và bạn’ll draw conclusions từ fake traffic.
Làm bạn ngay cả cần này?
Honestly? Probably không, nếu bạn chạy một nhỏ site. Log file analysis pays off cho lớn các trang — tens of thousands of URLs, ecommerce với lots of filtered các trang, các trang going qua một migration, hoặc các trang nơi Google says điều này “discovered” các trang nhưng không bao giờ được lập chỉ mục them. Nếu trang web của bạn có vài hundred các trang và they all nhận được crawl fine, của bạn time là tốt hơn spent elsewhere. (Giống nhau logic as crawl budget — hầu hết các trang không cần để worry về điều này.)
Muốn đó real workflow — verifying bots properly, finding crawl waste, spotting orphan các trang? Chuyển để đó Advanced tab.
Evidence for this claim Web-server access logs record HTTP requests and commonly include request, response-status, user-agent, and timing fields depending on configuration. Scope: Apache HTTP Server access-log behavior; other servers vary by configuration. Confidence: high · Verified: Apache HTTP Server: Log Files Evidence for this claim User-agent text alone does not authenticate Googlebot; Google recommends DNS verification or matching published IP ranges. Scope: Google crawler verification, applicable when classifying log traffic. Confidence: high · Verified: Google Search Central: Verify GooglebotTL;DR — Logs là đó unsampled ground truth cho crawling: mỗi yêu cầu, mỗi bot, mỗi mã trạng thái. Đó non-negotiable đầu tiên step là verifying Googlebot/Bingbot qua reverse + forward DNS (hoặc Google published IP-range JSON) — người dùng agents là spoofed constantly, và bạn làm all crawl math on đó verified set chỉ. Thì đọc đó logs cho hầu hết/least-được crawl URLs và sections, crawl frequency theo thời gian, các mã trạng thái prioritized by frequency, crawl waste (params, facets, internal tìm kiếm, infinite pagination), orphan và uncrawled các trang (cross-referenced so với một crawl), và đó mobile-so với-desktop Googlebot split. Bing publishes không chính thức IP ranges, so DNS để
*.search.msn.comlà đó phương thức ở đó. Logs complement GSC Crawl Số liệu — they không replace điều này. Và trong 2026, AI bots là hiện tại một huge share of điều gì cho thấy lên.
Vì sao logs là đó ground truth
Có three ways để “see” cách các công cụ tìm kiếm crawl bạn, và they không phải equal:
- MỘT crawl tool (Screaming Frog SEO Spider, Ahrefs Site Audit) simulates một crawl. Điều này tells bạn điều gì một bot có thể tìm, không điều gì Google đã làm fetch.
- GSC Crawl Số liệu summarizes đó real điều, nhưng đây là sampled, được tổng hợp, và capped (khoảng 1 000 các hàng, ~90 days, không theo-URL export).
- Máy chủ logs record đó real điều — mỗi yêu cầu, cho mỗi bot, với đó chính xác URL, timestamp, và mã trạng thái.
Đó Ahrefs log-file hướng dẫn I reviewed diễn đạt điều này plainly: máy chủ logs là “the most trustworthy source of information to understand the URLs that search engines have crawled.” (bản dịch) «đó hầu hết trustworthy nguồn of information để understand đó URLs đó các công cụ tìm kiếm có được crawl.» đó là đó toàn bộ reason này technique tồn tại. Khi I muốn để know điều gì Googlebot thực ra đã làm — không điều gì điều này có thể làm, không một rounded summary — I go để đó logs.
MỘT typical log line carries đó IP address, người dùng agent, URL path, timestamp, phương thức yêu cầu (GET/POST), và HTTP mã trạng thái. Mọi thứ dưới là chỉ slicing những các trường intelligently.
Khi bạn thực ra cần điều này (và khi bạn không)
Là honest với yourself ở đây. Log file analysis là một lớn-site tool. Điều này earns của nó
giữ on các trang với tens of thousands of URLs, ecommerce và faceted navigation,
các trang mid-migration, và các trang stuck trong Discovered – currently not indexed. As I
wrote trong my ngân sách crawl hướng dẫn, “Most sites don’t need to worry about crawl
budget, but there are few cases where you may want to take a look.” (bản dịch) «Hầu hết các trang không cần để worry về ngân sách crawl, nhưng có một vài cases nơi bạn có thể muốn để take một look.» Daniel
Waisberg of Google có đã làm một similar point về Crawl Số liệu — theo Công cụ tìm kiếm
Journal coverage, đó báo cáo không nhiều of một concern cho các trang dưới ~1 000
các trang.
Nếu của bạn một vài-hundred-trang site là đang được crawl fine, skip này và go cách sửa điều gì đó với hơn leverage.
Cách nhận của bạn logs (đó hardest part là access)
Logs trực tiếp wherever đó yêu cầu thực ra terminated:
- Apache & Nginx → đó Apache “combined” log format (đó hầu hết phổ biến).
- Microsoft IIS → W3C format.
- AWS ELB/ALB → ELB format.
- CDNs (Cloudflare, Fastly, Akamai) → của họ own log exports. Này matters: tại một CDN-fronted site, an origin-chỉ log misses edge-được lưu đệm hits, so pull logs tại đó layer đó bot thực ra reached.
Aim cho 30 days minimum, 90 ideal, để capture crawl-frequency biến động. Và plan cho friction — getting access để máy chủ logs là thường đó genuinely hard part (DevOps gatekeeping). Ngay cả Googlers, trong một migration episode of Tìm kiếm Off đó Record, flagged cách hard log files có thể là để obtain trong thực tế. Budget time cho đó yêu cầu.
Logs không chỉ bot traffic — they capture mỗi yêu cầu, including real khách truy cập, và có thể carry query-string các giá trị, session identifiers, hoặc other sensitive dữ liệu alongside đó URL path. OWASP logging hướng dẫn là blunt về này: authentication credentials, access tokens, và personally identifiable information generally không nên land trực tiếp trong một log; they nên là đã xóa, masked, hoặc hashed đầu tiên. Xây dựng đó vào của bạn access controls và export xử lý trước khi bạn hand một log file để anyone cho analysis, không sau.
Step 1 — Verify đó bots là real (làm này trước bất cứ điều gì khác)
Này là đó step hầu hết các hướng dẫn wave tại trong một line. Không. Nhiều bots pretend để là Googlebot để nhận past firewalls (Ahrefs). Người dùng agent là unauthenticated text; treat mỗi “Googlebot” line as một claim để là proven.
Googlebot — hai hợp lệ các phương thức:
- Reverse + forward DNS (đó bidirectional kiểm tra). Google own steps: chạy một
reverse DNS lookup on đó IP từ của bạn logs với đó
hostcommand; verify đó domain làgooglebot.com,google.com, hoặcgoogleusercontent.com; thì chạy một forward DNS lookup on đó hostname và verify điều này resolves lại để đó original IP. Đó forward step là điều gì làm này trustworthy — một spoofer có thể point reverse DNS tại một*.googlebot.comname, nhưng chỉ đó round-trip lại để đó giống nhau IP proves điều này. (Commands cho macOS/Linux và Windows là trong đó Scripts tab.) - Match so với Google published IP ranges. Google publishes JSON files of
của nó crawler IPs trong CIDR format —
common-crawlers.jsoncho Googlebot và friends, plusspecial-crawlers.json, người dùng-triggered-fetcher files, và an all-Googlegoog.json. As I noted trong my Googlebot hướng dẫn, Google “provided a list of public IPs you can use to verify the requests are from Google… You can compare this to the data in your server logs.” (bản dịch) «provided một list of công khai IPs bạn có thể dùng để verify đó các yêu cầu là từ Google… Bạn có thể so sánh này để đó dữ liệu trong máy chủ của bạn logs.»
Bingbot — DNS chỉ. Này là đó key contrast: Bing không officially
publish IP ranges. Bing own wording là đó “…like other search engines,
Bing does not publish a list of IP addresses or ranges from which we crawl the
Internet,” (bản dịch) «…như other các công cụ tìm kiếm, Bing không publish một list of IP addresses hoặc ranges từ mà we crawl đó Internet,» vì “the IP addresses or ranges we use can change any time.” (bản dịch) «đó IP addresses hoặc ranges we dùng có thể thay đổi bất kỳ time.» So
cho Bingbot bạn làm reverse + forward DNS để một hostname ending trong
*.search.msn.com (e.g. msnbot-157-55-33-18.search.msn.com), hoặc dùng đó
Verify Bingbot tool. (Microsoft
có since đã phát hành một bingbot IP JSON, nhưng của nó chính thức verification hướng dẫn
vẫn centers on DNS precisely vì IPs thay đổi.)
Thì throw out đó fakes. Làm all of của bạn crawl math on đó verified set chỉ. Unverified “Googlebot” là gần như luôn một scraper hoặc một spoofed bot và belongs trong một security review, không của bạn crawl-waste analysis.
Step 2 — Điều cần tìm
Khi bạn là hoạt động với verified hits, ở đây đó đọc:
- Hầu hết & least được crawl URLs và sections. Xếp hạng các yêu cầu by URL và by directory. Này là nơi của bạn ngân sách crawl là thực ra going — và đây là thường surprising.
- Crawl frequency theo thời gian. Trend crawl by URL/section để catch drops (một migration broke điều gì đó) hoặc spikes (một new section, hoặc một spider trap spinning lên infinite URLs).
- Các mã trạng thái bots hit, prioritized by frequency. Quantify
200so với.301/302(và chains),404, và5xx. MỘT404hit 5 000×/week là một khác nhau vấn đề hơn một404hit khi — cách sửa by crawl frequency, không by mere existence. - Crawl waste. Faceted nav, URL parameters, internal kết quả tìm kiếm, và infinite calendars/pagination có thể eat một lớn share of ngân sách crawl on bad offenders. Logs cho thấy chính xác mà junk patterns đó bots là burning time on.
- Orphan & uncrawled các trang. Này cần cả hai datasets. Cross-reference logs so với một site crawl: URLs trong đó logs nhưng không trong đó crawl = orphans, old các chuyển hướng, hoặc externally-linked các trang; URLs trong đó crawl nhưng không trong đó logs = các trang Google có không bao giờ fetched.
- Mobile so với. desktop Googlebot. Split by người dùng agent. Post mobile-đầu tiên, điều này nên là majority Googlebot Smartphone — một desktop-nặng split là worth một look.
- Thời gian phản hồi & crawl health. Rising average thời gian phản hồi correlates với reduced crawling. Theo SEJ ghi-lên of Waisberg hướng dẫn: “Watch out for a consistent increase in average response time. Google says it might not affect crawl rate immediately, but it’s a good indicator that your servers might not be handling all the load.” (bản dịch) «Watch out cho một consistent increase trong average thời gian phản hồi. Google says điều này có thể không ảnh hưởng tốc độ crawl immediately, nhưng đây là một good indicator đó của bạn các máy chủ có thể không là xử lý all đó load.»
Điều gì logs không tell bạn
Giữ những straight hoặc bạn’ll over-đọc đó dữ liệu:
- Crawl ≠ chỉ mục. MỘT URL Googlebot fetches daily có thể stay unindexed indefinitely. Logs prove fetching, không chỉ mục status — pair them với GSC Trang Lập chỉ mục / URL Inspection để learn đó chỉ mục side.
- Crawl ≠ xếp hạng, và hơn crawling không help. As I’ve đã nói repeatedly, “The rate of crawling isn’t going to impact your rankings.” (bản dịch) «Đó rate of crawling không going để impact của bạn thứ hạng.» không chase crawl volume as nếu điều này đã là một xếp hạng lever.
noindexkhông reduce crawling.noindexcontrols indexation, không crawling — để thực ra dừng đó crawl bạn dùng robots.txt hoặc một mã trạng thái.- Crawl ≠ model training hoặc citation. MỘT verified hit từ GPTBot, ClaudeBot, hoặc PerplexityBot proves đó yêu cầu happened — một fetch tại đó layer. Điều này không prove đó trang đã là được dùng để train một model, retained anywhere downstream, hoặc cited trong một chat câu trả lời. Những là tách biệt, unobserved outcomes; không stretch một verified log line further hơn điều này goes.
Đó 2026 wrinkle: AI bots là all over của bạn logs hiện tại
Đó cast of characters trong một modern log file có changed. Trong my analysis of Cloudflare Radar dữ liệu (Đáp ứng đó New Web Các crawler), tìm kiếm-engine bots vẫn crawl đó hầu hết — nhưng AI bots là firmly trong second place và on track để overtake them trong một couple of năm. GPTBot, ClaudeBot, PerplexityBot, và friends hiện tại cho thấy lên heavily. Khi bạn segment của bạn verified hits by người dùng agent, không là surprised để tìm AI các crawler rivaling đó tìm kiếm engines cho share of các yêu cầu. (Screaming Frog’s Log File Analyser có đã thêm một dedicated AI-bot tutorial cho chính xác này.)
Cách này fits với đó rest of crawling
Logs là đó diagnostic layer dưới đó toàn bộ crawling cluster. họ là cách
bạn thực ra đo lường đó ngân sách crawl spend đó engines mô tả trong đó
abstract (Gary Illyes defines điều này as “the number of URLs Googlebot can and is
willing or is instructed to crawl” (bản dịch) «đó number of URLs Googlebot có thể và là willing hoặc là instructed để crawl»). họ là cũng cách bạn catch spider traps
red-handed — an infinite URL space từ một calendar hoặc facet cho thấy lên as một flood of
near-giống hệt các yêu cầu — và cách bạn xác nhận liệu của bạn crawl frequency
hoạt động (chính xác lastmod, liên kết nội bộ để quan trọng các trang) thực ra changed bot
behavior. Và remember they complement, không replace, GSC Crawl Số liệu: Crawl
Số liệu là đó sampled on-ramp; logs là đó unsampled, multi-bot, theo-URL detail.
AI summary
MỘT condensed take on đó Advanced version:
- Logs = ground truth. Crawl tools simulate, GSC samples; máy chủ access logs record mỗi yêu cầu, mỗi bot, với URL, timestamp, và mã trạng thái.
- Verify trước khi bạn analyze. Người dùng agents là spoofed constantly. Xác nhận
Googlebot qua reverse + forward DNS hoặc Google published IP-range JSON; xác nhận
Bingbot qua reverse DNS để
*.search.msn.com(Bing publishes không chính thức IP ranges). Làm all crawl math on đó verified set chỉ. - Điều cần đọc: hầu hết/least-được crawl URLs và sections; crawl frequency over time; các mã trạng thái prioritized by frequency (một 404 hit 5 000×/week ≠ khi); crawl waste (params, facets, internal tìm kiếm, infinite pagination); orphan và uncrawled các trang (cross-reference một crawl); mobile-so với-desktop Googlebot split; rising thời gian phản hồi as một crawl-health warning.
- không over-đọc điều này: crawl ≠ chỉ mục ≠ xếp hạng, hơn crawling không help
thứ hạng,
noindexkhông reduce crawling, và một verified AI-bot hit proves một fetch — không đó trang đã là dùng cho model training hoặc cited trong an câu trả lời. - Phạm vi: một lớn-site / ecommerce / migration tool. Hầu hết nhỏ các trang không cần điều này.
- 2026: AI bots (GPTBot, ClaudeBot, PerplexityBot) là hiện tại một major và growing share of log traffic.
- Complements GSC Crawl Số liệu — dùng cả hai.
Tài liệu chính thức
Chính-nguồn tài liệu cho verifying các crawler và reading crawl dữ liệu.
- Verify Các yêu cầu từ Google Các crawler và Fetchers — đó hai chính thức các phương thức: manual reverse + forward DNS, và matching so với đó published IP ranges.
- Cách verify Googlebot (Tìm kiếm Central Blog) — đó older companion post, vẫn cited.
- phổ biến-các crawler.json — Googlebot và phổ biến crawler IPs trong CIDR format (note: “the IP addresses in the JSON files are represented in CIDR format” (bản dịch) «đó IP addresses trong đó JSON files là represented trong CIDR format»). Companions: special-các crawler.json, người dùng-triggered-fetchers.json, và đó all-Google goog.json.
- Overview of Google các crawler và fetchers — mỗi Google người dùng agent bạn’ll see trong logs.
- Crawl Số liệu báo cáo (Help) và đó launch blog — đó sampled, chính thức view đó logs complement.
Bing / Microsoft
- Cách Verify Bingbot (Quản trị viên web help) — đó chính thức verification trang.
- Cách Verify đó Bingbot là Bingbot (blog) — đó canonical reverse + forward DNS phương thức, và đó statement đó Bing không publish IP ranges.
- Verify Bingbot tool — paste an IP để kiểm tra điều này.
Quotes từ đó nguồn
On-đó-record statements. Nơi supported, mỗi link là một deep link đó jumps để đó quoted passage on đó trang nguồn.
Google — verifying các crawler
- “Run a reverse DNS lookup on the accessing IP address from your logs, using the
hostcommand.” (bản dịch) «Chạy một reverse DNS lookup on đó accessing IP address từ của bạn logs, dùng đóhostcommand.» — Google Search Central tài liệu. Nhảy đến trích dẫn - “Verify that the domain name is either
googlebot.com,google.com, orgoogleusercontent.com.” (bản dịch) «Verify đó domain name là eithergooglebot.com,google.com, hoặcgoogleusercontent.com.» Nhảy đến trích dẫn - “Run a forward DNS lookup on the domain name retrieved in step 1 using the
hostcommand on the retrieved domain name.” (bản dịch) «Chạy một forward DNS lookup on đó domain name retrieved trong step 1 dùng đóhostcommand on đó retrieved domain name.» Nhảy đến trích dẫn - “Verify that it’s the same as the original accessing IP address from your logs.” (bản dịch) «Verify đó đây là đó giống nhau as đó original accessing IP address từ của bạn logs.» Nhảy đến trích dẫn
Bing — verification và không published IP ranges
- “Perform a reverse DNS lookup using the IP address from the logs to verify that it resolves to a name that end with search.msn.com.” (bản dịch) «Perform một reverse DNS lookup dùng đó IP address từ đó logs để verify đó điều này resolves để một name đó end với tìm kiếm.msn.com.» — Bing Quản trị viên web Blog. Nhảy đến trích dẫn
- “…like other search engines, Bing does not publish a list of IP addresses or ranges from which we crawl the Internet.” (bản dịch) «…như other các công cụ tìm kiếm, Bing không publish một list of IP addresses hoặc ranges từ mà we crawl đó Internet.» — và đó reason: “the IP addresses or ranges we use can change any time, so responding to requests differently based on a hardcoded list is not a recommended approach.” (bản dịch) «đó IP addresses hoặc ranges we dùng có thể thay đổi bất kỳ time, so responding để các yêu cầu differently dựa trên một hardcoded list không phải một được khuyến nghị approach.» Nhảy đến trích dẫn
Patrick Stox — on điều gì logs là cho (từ my hoạt động tại Ahrefs)
- Máy chủ logs là “the most trustworthy source of information to understand the URLs that search engines have crawled.” (bản dịch) «đó hầu hết trustworthy nguồn of information để understand đó URLs đó các công cụ tìm kiếm có được crawl.» Nhảy đến trích dẫn
- “Many bots pretend to be Googlebot to get past firewalls.” (bản dịch) «Nhiều bots pretend để là Googlebot để nhận past firewalls.» Nhảy đến trích dẫn
- “If you want to see hits from all bots and users, you’ll need access to your log files.” (bản dịch) «Nếu bạn muốn để see hits từ all bots và người dùng, bạn’ll cần access để của bạn log files.» Nhảy đến trích dẫn
- “The rate of crawling isn’t going to impact your rankings.” (bản dịch) «Đó rate of crawling không going để impact của bạn thứ hạng.» Nhảy đến trích dẫn
Gary Illyes, Google — ngân sách crawl (điều gì logs let bạn đo lường)
- Ngân sách crawl là “the number of URLs Googlebot can and is willing or is instructed to crawl.” (bản dịch) «đó number of URLs Googlebot có thể và là willing hoặc là instructed để crawl.» Đọc bài đưa tin
Daniel Waisberg, Google — thời gian phản hồi as một crawl-health tín hiệu (SEJ cách diễn đạt of his Crawl Số liệu hướng dẫn)
- “Watch out for a consistent increase in average response time. Google says it might not affect crawl rate immediately, but it’s a good indicator that your servers might not be handling all the load.” (bản dịch) «Watch out cho một consistent increase trong average thời gian phản hồi. Google says điều này có thể không ảnh hưởng tốc độ crawl immediately, nhưng đây là một good indicator đó của bạn các máy chủ có thể không là xử lý all đó load.» Nhảy đến trích dẫn
Verify một bot là thực sự Googlebot (reverse + forward DNS)
Người dùng agent trong một log line là chỉ text — scrapers spoof Googlebot để nhận past
firewalls. Đó chỉ trustworthy kiểm tra là bidirectional DNS: reverse-lookup đó IP,
xác nhận đó hostname là một Google domain, thì forward-lookup đó hostname và
xác nhận điều này resolves lại để đó giống nhau IP.
macOS / Linux (dùng host)
# 1) Reverse DNS the IP from your logs — it must end in googlebot.com,
# google.com, or googleusercontent.com
host 66.249.66.1
# → 1.66.249.66.in-addr.arpa domain name pointer crawl-66-249-66-1.googlebot.com
# 2) Forward DNS that hostname back — it must resolve to the same IP
host crawl-66-249-66-1.googlebot.com
# → crawl-66-249-66-1.googlebot.com has address 66.249.66.1Windows (dùng nslookup)
:: 1) Reverse DNS the IP — confirm it ends in a Google domain
nslookup 66.249.66.1
:: 2) Forward DNS the returned hostname — confirm it matches the original IP
nslookup crawl-66-249-66-1.googlebot.comNếu đó reverse lookup không land on một Google domain, hoặc đó forward lookup
không trả về đó original IP, điều này là không Googlebot. (Cho Bingbot, chạy đó
chính xác giống nhau hai steps nhưng expect một hostname ending trong search.msn.com.) Bạn có thể
cũng skip DNS và match đó IP so với Google published ranges
(phổ biến-các crawler.json,
trong CIDR format).
Extract và count crawler hits từ một thô log
Nhanh một-liners cho một combined-format Apache/Nginx access log. (Những select on đó người dùng-agent string — remember để verify đó IPs trước trusting đó được tính.)
# Pull only the lines claiming to be Googlebot
grep -i "googlebot" access.log > googlebot-hits.log
# Count Googlebot hits per URL, most-crawled first
# (combined format: $7 is the request path)
grep -i "googlebot" access.log \
| awk '{print $7}' \
| sort | uniq -c | sort -rn | head -50
# Count Googlebot hits per status code (combined format: $9 is the status)
grep -i "googlebot" access.log \
| awk '{print $9}' \
| sort | uniq -c | sort -rn
# List the URLs Googlebot hit that returned a 404, by frequency
grep -i "googlebot" access.log \
| awk '$9 == 404 {print $7}' \
| sort | uniq -c | sort -rn
# Get the unique IPs claiming Googlebot — the list you then verify by DNS
grep -i "googlebot" access.log | awk '{print $1}' | sort -uAdjust đó $7/$9 trường positions nếu của bạn log format differs (IIS/W3C và ELB
order các trường differently).
Log file analysis workflow
- Nhận đó logs — máy chủ access logs, hoặc CDN/load-balancer exports. Tại một CDN-fronted site, pull edge logs cũng (origin misses bộ nhớ đệm hits).
- Grab đủ window — 30 days minimum, 90 ideal.
- Verify đó bots đầu tiên — reverse + forward DNS, hoặc Google published
IP ranges. Bingbot → DNS để
*.search.msn.com. - Drop đó fakes — làm all crawl math on đó verified set; route spoofed “Googlebot” để một security review.
- Xếp hạng hầu hết/least-được crawl URLs và sections — see nơi budget goes.
- Trend crawl frequency theo thời gian — catch drops (migrations) và spikes (new sections, spider traps).
- Tally các mã trạng thái —
200/301-302chains /404/5xx, và prioritize các cách sửa by crawl frequency, không mere existence. - Hunt crawl waste — parameters, faceted nav, internal tìm kiếm, infinite pagination/calendars.
- Tìm orphans & uncrawled các trang — cross-reference logs so với một site crawl (trong logs không crawl = orphan; trong crawl không logs = không bao giờ fetched).
- Kiểm tra đó mobile so với. desktop Googlebot split — nên là Smartphone-majority.
- Watch average thời gian phản hồi — một rising trend có thể throttle crawling.
- Segment AI bots — GPTBot, ClaudeBot, PerplexityBot, etc. là hiện tại một real share of traffic.
- Xác nhận với GSC — pair crawl findings với Crawl Số liệu và URL Inspection (crawl ≠ chỉ mục).
Điều cần xây dựng trong của bạn analysis sheet
Liệu bạn dùng Screaming Frog’s Log File Analyser, BigQuery, hoặc đó Ahrefs log-file template, bạn là building đó giống nhau handful of views. Sau verifying bots, parse mỗi line vào các cột và pivot.
Các cột để parse out of mỗi log line
| Cột | Từ đó log line | Vì sao điều đó quan trọng |
|---|---|---|
| IP address | $1 (combined format) | Đó điều bạn verify by DNS / IP range |
| Verified? | derived | Filter mỗi pivot để verified chỉ |
| Bot / người dùng agent | UA string | Segment Googlebot Smartphone so với. Desktop so với. AI bots |
| URL path | $7 | Group crawl by URL và by directory |
| Section / directory | derived từ path | Roll lên crawl spend by site area |
| Timestamp | [date] trường | Trend crawl frequency theo thời gian |
| Phương thức | GET/POST | Spot odd yêu cầu patterns |
| Mã trạng thái | $9 | 200 / 3xx / 404 / 5xx breakdown |
Pivots để xây dựng
- Crawl theo URL (và theo directory), descending — hầu hết & least được crawl.
- Crawl theo mã trạng thái, thì mã trạng thái × URL (so một cao-frequency
404jumps out). - Crawl theo day/week, segmented by section — đó frequency trend.
- Crawl theo bot/người dùng agent — đó mobile/desktop split, plus AI-bot share.
- MỘT logs so với. crawl join (VLOOKUP/hợp nhất so với một Screaming Frog hoặc Ahrefs crawl export) để surface orphans và không bao giờ-fetched các trang.
Đó Ahrefs hướng dẫn ships một downloadable template đó sets những lên; đó log-file-analysis hướng dẫn links điều này.
Cách đọc một log file: điều cần tìm
MỘT repeatable lens cho bất kỳ log analysis. Verify đầu tiên; thì chạy những truyền.
1. Verify, thì analyze. Đó dữ liệu là chỉ as good as đó bot identity behind điều này. Reverse + forward DNS (hoặc IP ranges); analyze đó verified set chỉ. Spoofed bots → security, không SEO.
2. Nơi là budget going? (hầu hết so với. least được crawl.) Xếp hạng by URL và by section. Đó goal là để tìm attention spent on đó sai places — và quan trọng các trang getting cũng little.
3. Điều gì là bots hitting? (các mã trạng thái, by frequency.)
200 là healthy; 3xx chains, 404, và 5xx là leaks. Triage by cách thường
đó bot hits mỗi, không by liệu đó lỗi merely tồn tại.
4. điều gì là đang wasted? (crawl waste.) Parameters, faceted nav, internal tìm kiếm, infinite pagination/calendars — đó classic budget sinks. Logs name đó chính xác offending patterns.
5. điều gì là bị thiếu? (orphans & uncrawled.) Cross-reference logs với một crawl. Trong logs nhưng không crawl = orphan/old chuyển hướng/liên kết bên ngoài. Trong crawl nhưng không logs = Google không bao giờ fetched điều này.
6. Ai crawling? (bot segmentation.) Mobile so với. desktop Googlebot (nên là Smartphone-majority), plus đó AI-bot share đó hiện tại rivals các công cụ tìm kiếm.
7. Là đó máy chủ healthy? (thời gian phản hồi.) Rising average thời gian phản hồi là an sớm warning đó crawling có thể nhận throttled.
Tools cho log file analysis
- Screaming Frog SEO Log File Analyser — đó workhorse desktop tool. Điều này auto-verifies tìm kiếm bots và flags spoofed IPs, và có Người dùng-Agent + Verification-Status filters cho granular bot analysis. Import một SEO Spider crawl và dùng đó “Not In URL Data” (bản dịch) «Không Trong URL Dữ liệu» filter để tìm orphans — “URLs which were discovered in your logs, but are not present in the crawl data imported.” (bản dịch) «URLs mà đã là discovered trong của bạn logs, nhưng không phải present trong đó crawl dữ liệu imported.» Điều này cũng hiện tại có một dedicated AI-bot monitoring tutorial. Tool
- BigQuery (và Splunk / ELK Stack / Logflare / logz.io) — cho thô, lớn-quy mô storage và querying khi log volume là cũng big cho một desktop tool.
- Ahrefs — cross-reference logs so với an Ahrefs Site Audit crawl để tìm orphans và xác nhận điều gì đó bots reached. (My home tool.)
- Semrush Log File Trình phân tích, OnCrawl, Botify, JetOctopus — other crawl+log các nền tảng, đó latter three aimed tại enterprise.
- GSC Crawl Số liệu — đó chính thức, free, sampled on-ramp. Complementary để thô logs, không một replacement: đây là được tổng hợp, capped, và có không theo-URL export.
- Verify Bingbot tool — paste an IP để xác nhận đây là thực sự Bingbot: bing.com/toolbox/verify-bingbot.
Mistakes để tránh
- Trusting đó “Googlebot” string không có verifying điều này. Vì sao đây là sai: đó
người dùng agent là unauthenticated text — anyone có thể put
Googlebottrong một yêu cầu header để slip past defenses hoặc pollute của bạn analysis với fake traffic. Làm thay vì: chạy reverse + forward DNS (hoặc match so với Google published IP ranges) trước một single line được tính toward của bạn crawl math. - Treating GSC Crawl Số liệu as đó đầy đủ picture. Vì sao đây là sai: đây là sampled, rounded, capped tại khoảng 1 000 các hàng và ~90 days, với không theo-URL export — điều này summarizes, điều này không record. Làm thay vì: dùng logs as đó ground truth và Crawl Số liệu as một complementary, nhanh hơn on-ramp.
- Pulling origin-chỉ logs on một CDN-fronted site. Vì sao đây là sai: an origin log misses mỗi yêu cầu đó CDN phân phối từ edge bộ nhớ đệm, so bạn là analyzing an incomplete picture of điều gì bots thực ra đã nhận. Làm thay vì: pull logs tại đó layer đó bot thực ra reached — đó CDN/edge export, không chỉ đó origin máy chủ.
- Sửa 404s và 5xx các lỗi by existence, không frequency. Vì sao đây là sai:
một
404hit khi một month và một404hit 5 000 times một week không phải đó giống nhau vấn đề, nhưng treating mỗi lỗi line as equally urgent wastes cách sửa effort. Làm thay vì: xếp hạng status-code các vấn đề by cách thường bots thực ra hit them. - Chasing crawl volume as nếu điều này đã là một xếp hạng lever. Vì sao đây là sai: hơn crawling không move thứ hạng — đây là một diagnostic tín hiệu, không một growth chỉ số. Làm thay vì: dùng crawl frequency để catch các vấn đề (drops, spider traps), không as một KPI để maximize.
- Dùng
noindexđể try để dừng crawling. Vì sao đây là sai:noindexcontrols indexation, không crawling — Google vẫn có để fetch đó trang để see đó tag. Làm thay vì: block đó crawl itself với robots.txt hoặc một status code nếu đó là đó thực tế goal. - Calling một URL “orphaned” từ logs alone. Vì sao đây là sai: một URL đó cho thấy lên trong logs nhưng không trong một fresh crawl có thể là an old chuyển hướng, an liên kết bên ngoài, hoặc một genuine orphan — logs alone không thể tell bạn mà. Làm thay vì: cross-reference logs so với an thực tế site crawl trước drawing conclusions either way.
Standing KPIs cho log file analysis
Verified-bot share
- Chỉ số: Verified Googlebot/Bingbot hits ÷ all hits claiming để là Googlebot/Bingbot by người dùng agent.
- Điều gì điều này tells bạn: Cách nhiều of của bạn “bot traffic” là thực ra spoofed scrapers thay vì real các công cụ tìm kiếm.
- Cách pull điều này: Chạy mỗi claimed-bot IP qua reverse + forward DNS (hoặc đó IP-range JSON) và count đó truyền rate.
- Benchmark / realistic range: Không universal number — phụ thuộc vào cách aggressively trang web của bạn là scraped. MỘT thấp hoặc dropping verified share là đó tín hiệu để act on, không một fixed ngưỡng.
- Cadence: Mỗi khi bạn pull một new log window.
Crawl waste share
- Chỉ số: % of verified bot các yêu cầu hitting parameters, facets, internal tìm kiếm, hoặc infinite pagination/calendars.
- Điều gì điều này tells bạn: Cách nhiều of của bạn ngân sách crawl là going để junk URLs thay vì các trang đó quan trọng.
- Cách pull điều này: Segment verified hits by URL pattern (query strings, known facet/tìm kiếm paths).
- Benchmark / realistic range: Không universal hình — genuinely situational, depending trên trang web của bạn URL structure và faceting. Establish của bạn own baseline on đó đầu tiên pull, thì track đó trend.
- Cadence: 30–90 day log window; re-kiểm tra sau bất kỳ cleanup (robots.txt rules, parameter xử lý, pagination các cách sửa).
Status-code mix, weighted by frequency
- Chỉ số:
200/3xx/404/5xxas một share of verified bot các yêu cầu. - Điều gì điều này tells bạn: Nơi bots là burning fetches on các lỗi thay vì trực tiếp nội dung, và liệu đó là getting tệ hơn.
- Cách pull điều này: Tally đó status-code trường từ verified log lines.
- Benchmark / realistic range: Không universal đích — phụ thuộc vào site age
và chuyển hướng history. Watch đó trend, không một single snapshot; một rising
404/5xxshare là đó actionable tín hiệu. - Cadence: 30–90 days, hoặc immediately sau một migration.
Mobile so với. desktop Googlebot split
- Chỉ số: Share of verified Googlebot hits từ đó Smartphone người dùng agent so với. Desktop.
- Điều gì điều này tells bạn: Liệu Google là thực ra crawling bạn mobile-đầu tiên, as dự kiến post mobile-đầu tiên lập chỉ mục.
- Cách pull điều này: Segment verified hits by đó Googlebot UA string (Smartphone so với. Desktop).
- Benchmark / realistic range: Nên là Smartphone-majority cho hầu hết các trang; một desktop-nặng split là worth investigating, không một hard failure on của nó own.
- Cadence: Mỗi log pull.
Orphan / uncrawled trang count
- Chỉ số: URLs trong logs nhưng không trong một fresh crawl (orphans/old links), và URLs trong một fresh crawl nhưng không trong logs (không bao giờ fetched).
- Điều gì điều này tells bạn: Các trang bots không thể easily reach, và các trang bạn là linking để đó Google có không bao giờ bothered để fetch.
- Cách pull điều này: Join đó verified-hit URL list so với một site crawl export (Screaming Frog, Ahrefs).
- Benchmark / realistic range: Hoàn toàn situational — phụ thuộc vào site size và cách recently bạn migrated hoặc restructured. Track đó count theo thời gian thay vì comparing để an bên ngoài number.
- Cadence: Quarterly đối với các trang web lớn, hoặc immediately sau một migration.
Ready-để-copy prompts cho log analysis
Những là cho interpreting log dữ liệu bạn đã pulled và verified — không cho generating log dữ liệu (không bao giờ let an AI invent log lines hoặc số liệu).
Summarize crawl waste từ một URL sample
Here is a list of URL paths that verified Googlebot hits landed on, one per
line, from my server logs. Group them into patterns (query parameters,
faceted navigation, internal search, pagination/calendars, or "looks like a
real page"), and tell me which pattern has the most URLs. Don't invent URLs
that aren't in the list — only group what I've pasted.
[paste your URL list here]Kết quả mong đợi: một nhỏ set of named clusters với được tính, plus một flag on mà cluster looks như đó biggest crawl-waste offender — treat điều này as một starting trỏ đến verify so với đó thực tế paths, không một cuối câu trả lời.
Prioritize một status-code breakdown
I have this table of HTTP status codes and how many times verified Googlebot
hit each one over the last 30 days. Rank them by which I should fix first,
weighting frequency over severity — a 404 hit 5,000 times matters more than a
500 hit twice. Explain the reasoning in one line per row.
status_code, hit_count
[paste your table here]Kết quả mong đợi: đó giống nhau các hàng re-ordered by cách sửa priority với một một-line reason mỗi — sanity-kiểm tra đó reasoning so với của bạn own site context trước acting.
Draft một log-access yêu cầu để DevOps
Write a short, plain-English email to my DevOps/hosting team asking for
30-90 days of raw web server access logs (Apache/Nginx combined format, or
our CDN's edge logs if we're behind one) for [site name]. Explain in one
sentence why I need it (verifying real Googlebot/Bingbot crawl activity vs
GSC's sampled report) and ask what export format and delivery method works
for them.Kết quả mong đợi: một ngắn draft email bạn có thể edit với của bạn thực tế site name và gửi — review điều này yourself trước sending, này tool sẽ không gửi messages on của bạn behalf.
Giải thích một DNS verification kết quả
I ran a reverse DNS lookup on an IP from my server logs and then a forward
DNS lookup on the hostname it returned. Here's the raw output from the
`host` command. Tell me plainly whether this confirms the request came from
real Googlebot or Bingbot, and point to exactly which line proves or
disproves it.
[paste your host/nslookup output here]Kết quả mong đợi: một đơn giản-language verdict tied để đó cụ thể line trong của bạn output — treat điều này as một second opinion, không một replacement cho knowing đó thực tế rule (hostname ends trong một Google/Bing domain, và đó forward lookup trả về đó original IP).
Các tài nguyên worth của bạn time
My related writing
- Cách Làm an SEO Log File Analysis [Template Được bao gồm] — đó Ahrefs cornerstone hướng dẫn I reviewed; đó cách diễn đạt và template này bài viết leans on.
- Khi Nên Bạn Worry Về Ngân sách crawl? — khi log analysis là (và không) worth của bạn time.
- Điều gì Là Googlebot & Cách Làm Điều này Hoạt động? — đó các crawler và IP-verification background.
- Đáp ứng đó New Web Các crawler: AI Bots Là Closing trong on Công cụ tìm kiếm Bots — vì sao của bạn logs look khác nhau trong 2026.
- Đó Beginner Hướng dẫn để SEO kỹ thuật — nơi crawling và logs fit trong đó bigger picture.
My speaking
- Cách Tìm kiếm Hoạt động (SlideShare) — my walkthrough of crawling: Googlebot as 1 000+ các hệ thống với nhiều specialized các crawler (Desktop, Mobile, Image, News, Video, Quảng cáo) sharing một crawl-budget pool, và các yêu cầu mostly originating out of Mountain View — một handy log sanity-kiểm tra alongside, không bao giờ thay vì, proper verification. (Standing disclaimer: “This is my understanding of systems… not going to be 100% complete or accurate.” (bản dịch) «Này là my understanding of các hệ thống… không going để là 100% hoàn tất hoặc chính xác.»)
Từ others
- Screaming Frog Log File Analyser — người dùng hướng dẫn và AI-bot tutorial.
- Search Engine Land — Log file analysis hướng dẫn (Kody Wirth) — practical walkthrough covering formats, tools, và điều gì patterns để tìm.
- Search Engine Journal — Cách Dùng Google Crawl Số liệu Báo cáo — Daniel Waisberg (Google) hướng dẫn on reading Crawl Số liệu, including đó phản hồi-time warning; đó best companion để log analysis.
- Search Engine Journal — Lập chỉ mục & Ngân sách crawl (Illyes + Splitt) — on-đó-record definition of ngân sách crawl và cách quality drives crawl demand.
- Công cụ tìm kiếm Roundtable — Bingbot IP addresses đã phát hành — covers đó nuance đó Microsoft eventually published một bingbot IP JSON mặc dù chính thức hướng dẫn vẫn centers on DNS.
- Conductor — Log File Analysis cho SEO — accessible explainer of điều gì logs reveal và cách act on them.
- r/TechSEO — đó community cho crawl/chỉ mục gỡ lỗi.
Videos
- Google Search Central (YouTube) — Martin Splitt crawling/kết xuất explainers và đó Cách Google Search Hoạt động series; hữu ích background on đó bots whose hits bạn là verifying trong của bạn logs. Channel
Tự kiểm tra
Five các câu hỏi on verifying bots và reading điều gì của bạn logs thực ra cho thấy.
Nhật ký thay đổi
Đã cập nhật 8 thg 8, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.
Đã cập nhật 30 thg 7, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.
Đã cập nhật 18 thg 7, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.