Hướng dẫn về YandexBot
Điều gì YandexBot là, cách spot và verify điều này, đó Yandex-chỉ Sạch-param directive, vì sao Crawl-delay là dead, cách điều này xử lý JavaScript, và cách điều này compares để Googlebot và Bingbot.
Ngôn ngữ
1 tín hiệu bằng chứng trên trang này
- Công cụ trực tuyến liên quanGooglebot Verifier
YandexBot là Yandex main web crawler — đó bot đó discovers và fetches các trang cho Yandex Tìm kiếm, đó engine với ~70%+ share trong Russia. Của nó robots.txt token là YandexBot (main lập chỉ mục bot chỉ) so với. Yandex (đó rộng hơn bot family). Điều này hỗ trợ một Yandex-chỉ directive, Sạch-param, đó consolidates URL parameters với không Google/Bing tương đương — và điều này đã dừng honoring Crawl-delay on February 22, 2018 (some SEO các hướng dẫn vẫn wrongly claim nếu không). JavaScript kết xuất là beta và 'tại đó bot discretion,' và đó 2023 nguồn-code leak suggested có không tách biệt JS kết xuất hệ thống đó way Google có một. Verify một real YandexBot by reverse-thì-forward DNS để một yandex.ru/.net/.com host — đó giống nhau technique Google và Bing dùng cho của họ own bots — không by người dùng-agent string.
Evidence for this claim Yandex documents its search robots and their user-agent identifiers in Yandex Webmaster Help. Scope: Current official Yandex robot list. Confidence: high · Verified: Yandex Webmaster: Yandex robots Evidence for this claim Yandex provides an official method for checking whether an IP address belongs to a Yandex robot; a user-agent string alone can be spoofed. Scope: Current Yandex robot verification guidance. Confidence: high · Verified: Yandex Webmaster: Verify a robotTóm tắt — YandexBot là crawler cho Yandex Tìm kiếm — giống nhau job Googlebot làm cho Google và Bingbot làm cho Bing. nó visits của bạn các trang, downloads them, và adds them để Yandex chỉ mục. Liệu bạn nên care về nó xuất hiện xuống để một câu hỏi: làm bạn có bất kỳ audience hoặc business trong markets nơi Yandex là relevant? nếu có, nó matters. nếu không, nó mostly chỉ traffic trong của bạn nhật ký.
Điều gì YandexBot là
Khi bạn see YandexBot trong của bạn máy chủ nhật ký, đó crawler cho Yandex —
công cụ tìm kiếm đó dominates tìm kiếm trong Russia way Google dominates phần lớn của
rest của world. chỉ như Googlebot và Bingbot, YandexBot follows links,
đọc sitemaps, downloads của bạn các trang, và hands them off để là đã thêm để Yandex
tìm kiếm chỉ mục.
Đó reason điều này nhận của nó own bài viết — thay vì “just block all the non-Google bots” (bản dịch) «chỉ block all đó non-Google bots» — là geography. Yandex là một rounding lỗi worldwide, nhưng trong Russia điều này holds khoảng 71% of đó tìm kiếm market versus Google ~27% (StatCounter, June 2026). So nếu bạn sell để hoặc serve mọi người trong Russia (và trong lịch sử some nearby markets), YandexBot là của bạn gateway để hầu hết of đó tìm kiếm traffic.
Cách spot nó
reliable identity kiểm tra cho YandexBot là DNS lookup, không người dùng-agent string — ở đây Vì sao, và Cách nó hoạt động.
YandexBot người dùng-agent contains word YandexBot. Yandex own tài liệu
lists đầy đủ string as:
Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/81.0.4044.268nhưng đó string alone không prove bất cứ điều gì: anyone có thể fake nó. Scrapers và bad bots routinely pretend để là YandexBot. So người dùng-agent là chỉ filter cho mà log lines là worth kiểm tra — thực tế proof là reverse-sau đó-forward DNS lookup confirming yêu cầu nghĩ ra từ thực Yandex host, giống nhau technique Google và Bing sử dụng cho của họ own các crawler (covered trong Nâng cao tab).
nên bạn block nó?
Đây là thực câu hỏi phần lớn mọi người là asking. nhanh way để think về nó:
- bạn có Russia/CIS-facing business? không block nó — bạn’d là cutting yourself out của dominant công cụ tìm kiếm ở đó.
- bạn có zero Russia-facing audience và crawling là straining của bạn
máy chủ? sau đó blocking hoặc slowing nó là reasonable call. Bạn có thể tell nó để
stay out với couple của lines trong của bạn
robots.txt.
Một caveat on đó block side: một robots.txt rule là một yêu cầu, không
enforcement. Yandex own tài liệu warn đó some of của nó robots có thể bỏ qua robots.txt
directives, so nếu bạn cần guaranteed exclusion — không chỉ “please don’t” (bản dịch) «please không» — block
by verified IP tại đó máy chủ/firewall cấp độ thay vì (see đó Advanced tab cho mà
Yandex bots này ảnh hưởng).
Một quan trọng gotcha, và đây là đó giống nhau trap as Google: putting một trang trong
robots.txt không xóa điều này từ Yandex kết quả tìm kiếm — điều này chỉ dừng
Yandex từ reading đó trang. Yandex says này trong của nó own tài liệu, và adds một
condition worth knowing: nếu bạn cũng block một trang trong robots.txt, Yandex “can’t
index them and detect your instructions” (bản dịch) «không thể chỉ mục them và detect của bạn instructions» — meaning một noindex tag chỉ hoạt động nếu
bạn let Yandex fetch đó trang để see điều này. Blocking crawling và thêm noindex on
đó giống nhau URL cancels đó noindex out. Để thực ra giữ một trang out, cho phép crawling
và thêm một noindex tag thay vì.
Muốn kỹ thuật version — chính xác robots.txt tokens, Yandex unique Sạch-param directive, Cách verify thực YandexBot, và Cách nó xử lý JavaScript? Chuyển để Nâng cao tab.
Evidence for this claim Yandex documents its search robots and their user-agent identifiers in Yandex Webmaster Help. Scope: Current official Yandex robot list. Confidence: high · Verified: Yandex Webmaster: Yandex robots Evidence for this claim Yandex provides an official method for checking whether an IP address belongs to a Yandex robot; a user-agent string alone can be spoofed. Scope: Current Yandex robot verification guidance. Confidence: high · Verified: Yandex Webmaster: Verify a robotTóm tắt — YandexBot là Yandex Tìm kiếm main lập chỉ mục crawler. Của nó
robots.txttokenYandexBottargets chỉ main lập chỉ mục bot;Yandextargets rộng hơn bot family. nó hỗ trợ Yandex-chỉ directive, Sạch-param, đó consolidates URL parameters — không Google/Bing tương đương — và nó đã dừng honoringCrawl-delayon February 22, 2018 (sử dụng Tốc độ crawl tool thay vì; some SEO các hướng dẫn vẫn wrongly claim Yandex hỗ trợ Crawl-delay). JavaScript kết xuất là performed tại crawler discretion, so ưu tiên SSR/pre-kết xuất cho cốt yếu nội dung.Disallow≠noindex(giống nhau trap as Google). Verify thực YandexBot by reverse-sau đó-forward DNS đểyandex.ru/yandex.net/yandex.comhost — giống nhau technique Google và Bing sử dụng — không bao giờ người dùng-agent string alone.
Điều gì YandexBot thực ra là
YandexBot là đó chính web crawler cho Yandex, đó Russian công cụ tìm kiếm. Điều này discovers URLs, fetches các trang, và feeds Yandex chỉ mục — đó giống nhau role Googlebot và Bingbot play cho của họ engines. Người dùng-agent string Yandex documents (on của nó “check that a robot belongs to Yandex” (bản dịch) «kiểm tra rằng một robot belongs để Yandex» trang) là:
Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/81.0.4044.268Yandex adds một hữu ích caveat tiếp theo để điều này: vì “the browser’s version may change,” (bản dịch) «đó trình duyệt version có thể thay đổi,»
điều này khuyến nghị không matching on một fixed Chrome version khi bạn là trying để
identify đó bot. Match đó YandexBot token, không Chrome/81.0.4044.268.
Crucially, “YandexBot” là thực sự chỉ đó main lập chỉ mục member of một family of
Yandex robots — YandexImages, YandexMetrika, YandexDirect, YandexMobileBot,
YandexAccessibilityBot, YandexRenderResourcesBot, YandexCalendar, và hơn — mỗi
independently controllable trong robots.txt. Nhiều các bài viết conflate “YandexBot” với
“all Yandex crawlers,” (bản dịch) «all Yandex các crawler,» mà là imprecise.
Một hơn worth knowing về: Yandex máy chủ-logs bảng hiện tại documents
YandexAdditionalBot (và một near-duplicate token, YandexAdditional) as một robot
đó “helps process robots.txt to prevent page content from appearing in Search
with Yandex AI responses,” (bản dịch) «helps xử lý robots.txt để ngăn trang nội dung từ appearing trong Tìm kiếm với Yandex AI các phản hồi,» applied để các trang đó main crawler có đã được lập chỉ mục.
Theo đó giống nhau bảng, điều này không take đó chung User-agent: * rules vào account
— so nếu bạn muốn để opt một trang out of Yandex AI features cụ thể, bạn cần
an rõ ràng User-agent: YandexAdditionalBot block, đó giống nhau pattern other
engines’ AI-crawler opt-outs dùng.
YandexBot so với. “Yandex” trong robots.txt — họ’re không giống nhau token
Đây là Yandex phần lớn non-obvious robots.txt quirk, và nó bites mọi người migrating
từ Google-centric kỹ thuật SEO. Bắt đầu từ Yandex own worked ví dụ — nó
làm phạm vi split rõ ràng, comments được bao gồm:
User-agent: YandexBot # will be used only by the main indexing bot
Disallow: /*id=
User-agent: Yandex # will be used by all Yandex bots
Disallow: /*sid= # except the main indexing bot
User-agent: * # will not be used by Yandex bots
Disallow: /cgi-binđọc theo nghĩa đen, đó ví dụ own comments là tài liệu:
User-agent: YandexBot— dùng chỉ by đó main lập chỉ mục bot.User-agent: Yandex— dùng by Yandex bots hơn broadly — nhưng, theo đó ví dụ own comment on đó second block, “except the main indexing bot.” (bản dịch) «except đó main lập chỉ mục bot.» Đó rộng hơn token không universal ngay cả trong đó Yandex family.
Hai điều follow từ Yandex rules ở đây. Đầu tiên, precedence: “If the User-agent: Yandex string is detected, the User-agent: * string is ignored.” (bản dịch) «Nếu đó string là detected, đó string là đã bỏ qua.» So một generic
User-agent: * block sẽ không apply để Yandex bots nếu bạn đã cũng được viết một Yandex
block. Second — và này là đó một đó surprises security-minded readers — Yandex
warns đó “Some Yandex robots may ignore directives in robots.txt, including
those for User-agent: Yandex.” (bản dịch) «Some Yandex robots có thể bỏ qua directives trong , including những cho .» Không mỗi Yandex bot là guaranteed để obey một
blanket rule, mà là một hơn reason máy chủ-cấp độ verification và blocking
quan trọng cho đầy đủ exclusion.
Verifying nó thực sự YandexBot
Vì người dùng-agent là spoofable, Yandex tells bạn để verify với DNS, chính xác như Google và Bing làm cho của họ own các crawler. Yandex: “Some robots can disguise themselves as Yandex robots by indicating the relevant User Agent. You can check the authenticity of a robot using a reverse DNS lookup.” (bản dịch) «Some robots có thể disguise themselves as Yandex robots by indicating đó relevant Người dùng Agent. Bạn có thể kiểm tra đó authenticity of một robot dùng một reverse DNS lookup.» Đó được ghi lại phương thức:
- “Determine the IP address of the user agent in question using your server logs.” (bản dịch) «Determine đó IP address of người dùng agent trong câu hỏi dùng máy chủ của bạn logs.»
- “Use a reverse DNS lookup of the IP address to determine the host domain name.” (bản dịch) «Dùng một reverse DNS lookup of đó IP address để determine đó host domain name.»
- “Check whether the host belongs to Yandex. All Yandex robots have names ending
in
yandex.ru,yandex.netoryandex.com.” (bản dịch) «Kiểm tra xem đó host belongs để Yandex. All Yandex robots có names ending trong , hoặc .» (Nếu đó host name có một khác nhau ending, điều này không Yandex.) - “Make sure that the name is correct. Use a forward DNS lookup to get the IP address corresponding to the host name. It should match the IP address used in the reverse DNS lookup.” (bản dịch) «Hãy bảo đảm đó name là correct. Dùng một forward DNS lookup để nhận đó IP address corresponding để đó host name. Điều này nên match đó IP address dùng trong đó reverse DNS lookup.»
Và đó fail condition, trong Yandex words: “If the IP addresses do not match, it means that the host name is fake.” (bản dịch) «Nếu đó IP addresses không match, điều này có nghĩa là đó host name là fake.» Yandex cũng mentions an chính thức “IP address check tool” (bản dịch) «IP address kiểm tra tool» as an alternative để đang chạy đó lookups by hand.
Đây là giống nhau forward-confirmed reverse-DNS (FCrDNS) pattern all three major
engines land on — Google verifies so với googlebot.com/google.com/
googleusercontent.com, Bing so với *.search.msn.com, và Yandex so với
yandex.ru/yandex.net/yandex.com. None của them treat published IP list as
trustworthy đủ on của nó own. commands là trong Scripts tab; domain
suffixes là chỉ điều đó thay đổi giữa engines. (cho Google và Bing
versions, see Googlebot và Bingbot siblings.)
Controlling YandexBot với robots.txt
Yandex recognizes familiar cốt lõi đặt của directives, mỗi được định nghĩa trong của nó own tài liệu:
- Người dùng-agent — “Indicates the robot to which the rules listed in
robots.txtapply.” (bản dịch) «Indicates đó robot để mà đó rules listed trong apply.» - Disallow — “Prohibits crawling of sections or individual pages of the site.” (bản dịch) «Prohibits crawling of sections hoặc riêng lẻ các trang of đó site.»
- Cho phép — “Allows indexing site sections or individual pages.” (bản dịch) «Cho phép lập chỉ mục site sections hoặc riêng lẻ các trang.»
- Sitemap — “Specifies the path to the
Sitemapfile that is posted on the site.” (bản dịch) «Specifies đó path để đó file đó là posted on đó site.» - Sạch-param — “Indicates to the robot that the page URL contains parameters (like UTM tags) that should be ignored when indexing it.” (bản dịch) «Indicates để đó robot đó trang URL contains parameters (như UTM tags) đó nên là đã bỏ qua khi lập chỉ mục điều này.» (Yandex-chỉ — see dưới.)
Vài file requirements worth knowing: đó file phải được “a TXT file named
” (bản dịch) «một TXT file named»robots”, robots.txt,” (bản dịch) «, ,» của nó size không được exceed 500 KB, và đó máy chủ phải
trả về an HTTP 200 OK status cho điều này để là đọc.
Disallow ≠ noindex — giống nhau trap as Google
Đó single hầu hết misunderstood robots.txt fact carries straight over để Yandex, và
Yandex trạng thái điều này plainly: “Pages restricted in robots.txt can participate in
Yandex search. To remove pages from search, specify the noindex directive in the
HTML code of the page or configure the HTTP header.” (bản dịch) «Các trang restricted trong có thể participate trong Yandex tìm kiếm. Để xóa các trang từ tìm kiếm, specify đó directive trong đó HTML code of đó trang hoặc configure đó HTTP header.» Nói cách khác, Disallow
controls crawling, không lập chỉ mục — một disallowed URL có thể vẫn cho thấy lên trong Yandex
kết quả. Này là đó giống nhau conceptual trap Google có (I’ve được viết điều này lên cho Google
trong Được lập chỉ mục, though Bị chặn bởi robots.txt),
và đó cách sửa là giống hệt: để thực ra xóa một trang, cho phép crawling và thêm
noindex. Yandex spells out chính xác vì sao trong đó giống nhau section: “Do not restrict
such pages in robots.txt, or the Yandex bot can’t index them and detect your
instructions.” (bản dịch) «Không restrict such các trang trong , hoặc đó Yandex bot không thể chỉ mục them và detect của bạn instructions.» An chỉ mục-control directive chỉ hoạt động nếu đó crawler có thể fetch đó
trang để see điều này — Disallow và noindex on đó giống nhau URL là một contradiction: đó
Disallow dừng Yandex từ bao giờ reading đó noindex tag, so đó trang vẫn giữ
chính xác nơi điều này đã là.
Sạch-param — Yandex unique parameter directive
Sạch-param là đó single hầu hết Yandex-cụ thể directive cho an audience được dùng để Google và Bing, và điều này có không Google hoặc Bing tương đương. Của nó purpose, theo Yandex: “The Yandex robot uses this directive to avoid reloading duplicate information. This improves the robot’s efficiently and reduces the server load.” (bản dịch) «Đó Yandex robot dùng này directive để tránh reloading duplicate information. Này improves đó robot efficiently và reduces đó máy chủ load.» (Đó “efficiently” là một genuine typo on Yandex trực tiếp trang — I’m quoting điều này as-là thay vì silently sửa điều này.)
Đó vấn đề điều này solves: “The new parameter that doesn’t affect the page content may result in duplicate pages that should not be included in the search.” (bản dịch) «Đó new parameter đó không ảnh hưởng đó trang nội dung có thể kết quả trong duplicate các trang đó không nên là được bao gồm trong đó tìm kiếm.» Đó syntax:
Clean-param: p0[&p1&p2&..&pn] [path]Yandex own worked ví dụ — three các URL đó differ chỉ by ref theo dõi
parameter:
www.example.com/some_dir/get_book.pl?ref=site_1&book_id=123
www.example.com/some_dir/get_book.pl?ref=site_2&book_id=123
www.example.com/some_dir/get_book.pl?ref=site_3&book_id=123…collapse để một canonical URL (www.example.com/some_dir/get_book.pl?book_id=123)
với single directive:
User-agent: Yandex
Clean-param: ref /some_dir/get_book.plHai details làm Sạch-param easy để nhận sai. Đầu tiên, “The Clean-param directive
does not require mandatory combination with the Disallow directive” (bản dịch) «Đó Sạch-param directive không yêu cầu mandatory combination với đó Disallow directive» — điều này stands on
của nó own; bạn không Disallow đó parameter URLs. Second, đây là intersectional:
theo Yandex điều này “is intersectional, so it can be specified anywhere in the file,
regardless of the location.” (bản dịch) «là intersectional, so điều này có thể là specified anywhere trong đó file, regardless of đó location.» Unlike Allow/Disallow, mà là anchored để một
path, Sạch-param là một global directive bạn có thể drop anywhere trong đó file.
Yandex cũng notes điều này có thể xử lý some parameters tự động: “Parameters for analytics and tracking that don’t affect the page content may be automatically removed by the search engine if the algorithms determine that those parameters are insignificant.” (bản dịch) «Parameters cho analytics và tracking đó không ảnh hưởng đó trang nội dung có thể là tự động đã xóa by đó công cụ tìm kiếm nếu đó algorithms determine đó những parameters là insignificant.» Nhưng relying on Sạch-param cho đó ones đó quan trọng để bạn là đó deterministic move. Này là đó Yandex analog để cách Google hiện tại leans on các tín hiệu canonicalization (canonical tag, liên kết nội bộ) since điều này deprecated của nó old URL Parameters tool — Yandex chỉ cho bạn an rõ ràng robots.txt directive nơi Google không.
Crawl-delay là dead (since Feb 2018)
Nếu bạn đã đọc đó Crawl-delay hoạt động trong Yandex robots.txt, đó information là
stale. Yandex own dedicated trang là unambiguous: “From February 22, 2018, Yandex
doesn’t take into account the Crawl-delay directive.” (bản dịch) «Từ February 22, 2018, Yandex không take vào account đó Crawl-delay directive.»
Này là worth flagging vì ít nhất một widely-đọc SEO tài nguyên vẫn says đó
opposite. Ahrefs’ robots.txt hướng dẫn (authored by Joshua Hardwick, không me) hiện tại
trạng thái “Google no longer supports this directive, but Bing and Yandex do.” (bản dịch) «Google không lâu hơn hỗ trợ này directive, nhưng Bing và Yandex làm.» On
đó Yandex half, đó là contradicted by Yandex own hiện tại tài liệu. I’d
treat Yandex dedicated, dated trang as có thẩm quyền ở đây — nhưng đó rộng hơn lesson
là đó hữu ích một: verify một crawl-behavior claim so với đó engine trực tiếp tài liệu
trước khi bạn trust một secondhand hướng dẫn, vì những details drift và ngay cả good
sources go stale. (Note đó contrast với Bing, mà làm vẫn honor
crawl-delay — một of đó real Bingbot/YandexBot divergences.)
replacement là Tốc độ crawl setting trong Yandex Quản trị viên web, mà lets bạn influence Cách fast YandexBot fetches của bạn trang web. (Yandex Crawl-delay trang là essentially một sentence plus pointer để đó setting.)
Cách YandexBot xử lý JavaScript
Yandex JavaScript kết xuất là explicitly labeled beta (β) by Yandex itself, và đó default behavior là “at the bot’s discretion” (bản dịch) «tại đó bot discretion» — đó bot “will independently determine whether to execute JavaScript code on the site’s pages.” (bản dịch) «sẽ independently determine liệu để execute JavaScript code on đó site các trang.» Khi điều này làm, điều này có thể “assess the quality and completeness of the content on the pages with and without JavaScript” (bản dịch) «assess đó quality và tính đầy đủ of đó nội dung on đó các trang với và không có JavaScript» và serve whichever version là có khả năng hơn hữu ích để đó khách truy cập.
có một có ý nghĩa tension worth surfacing. Đó 2023 Yandex nguồn-code leak (exposed internal engineering tài liệu, không an chính thức statement — I’ll cover điều này hơn dưới) suggested một simpler picture. As Mike King wrote trong his Search Engine Land analysis of đó leak: “Yandex has no separate rendering system for JavaScript. They say this in their documentation and, although they have Webdriver-based system for visual regression testing called Gemini, they limit themselves to text-based crawl.” (bản dịch) «Yandex có không tách biệt kết xuất hệ thống cho JavaScript. They chẳng hạn này trong của họ tài liệu và, although they có Webdriver-based hệ thống cho visual regression kiểm thử called Gemini, they limit themselves để text-based crawl.» (Đó internal “Gemini” là một Yandex visual-regression kiểm thử tool — completely unrelated để Google Gemini AI model, despite đó shared name. Worth disambiguating so không ai conflates đó hai.) Và Dan Taylor Search Engine Journal ghi-lên concluded có “nothing new to suggest Yandex can crawl JavaScript yet outside of already publicly documented processes.” (bản dịch) «không có gì new để suggest Yandex có thể crawl JavaScript tuy vậy bên ngoài of đã publicly được ghi lại xử lý.»
So đó honest câu trả lời là neither “YandexBot renders JS just like Googlebot” (bản dịch) «YandexBot renders JS chỉ như Googlebot» nor “YandexBot never touches JavaScript.” (bản dịch) «YandexBot không bao giờ touches JavaScript.» đây là selective và beta, by Yandex own mô tả. Đó leak “no separate rendering system” (bản dịch) «không tách biệt kết xuất hệ thống» account là consistent với đó cách diễn đạt, nhưng đây là leaked internal material relayed qua ngành coverage, không một Yandex disclosure — so treat “architecturally simpler than Google” (bản dịch) «architecturally simpler hơn Google» as một plausible reading of hai consistent các tín hiệu, không một được ghi lại fact. không take either account on faith cho một route đó matters để bạn; kiểm thử điều này trực tiếp:
- Được kết xuất output: fetch đó trang với JavaScript disabled và so sánh điều này để đó JS-được kết xuất version. Nếu đó hai differ meaningfully, không assume Yandex saw đó được kết xuất một.
- Tài nguyên access: xác nhận đó JS, CSS, và API endpoints đó trang phụ thuộc vào
không blocked trong
robots.txt. YandexYandexRenderResourcesBotfetches render-time các tài nguyên, nhưng (theo Yandex own tài liệu of điều này) chỉ cho các trang đó main lập chỉ mục bot có thể đã reach — một blocked tài nguyên on an được phép trang vẫn sẽ không load. - Delayed và interactive nội dung: bất cứ điều gì đó loads sau đó
DOMContentLoadedevent hoặc behind một nhấp không guaranteed để render. Yandex advanced kết xuất settings (window.YandexRotorSettings) exist cụ thể cho các trang nơi “content loads with a delay” (bản dịch) «nội dung loads với một delay» — đó là một tín hiệu worth reading, không chỉ một config option.
Đó practical takeaway: Yandex own tài liệu khuyến nghị bạn “Prohibit rendering if SSR (Server-Side Rendering) or pre-rendering is implemented on the site,” (bản dịch) «Prohibit kết xuất nếu SSR (Máy chủ-Side Kết xuất) hoặc pre-kết xuất là implemented on đó site,» và note đó “Executing JavaScript code may create additional load on your server.” (bản dịch) «Executing JavaScript code có thể tạo additional load on máy chủ của bạn.» Nếu bạn muốn reliable Yandex lập chỉ mục of JS-nặng nội dung, serve điều này máy chủ-side rather hơn kiểm thử của bạn luck on client-side kết xuất.
Cho AJAX-style các trang, Yandex says “When indexing an AJAX site, the Yandex bot scans
the original URLs and executes JavaScript code on them” (bản dịch) «Khi lập chỉ mục an AJAX site, đó Yandex bot scans đó original URLs và executes JavaScript code on them» — và điều này có moved away
từ đó old HTML-snapshot hack: nếu bạn vẫn dùng đó deprecated
meta name="fragment" approach, “the bot will ignore it and index the original page.” (bản dịch) «đó bot sẽ bỏ qua điều này và chỉ mục đó original trang.»
Của nó modern khuyến nghị mirrors Google: “If the links on AJAX pages use the #
character, change the addresses to URLs without this character. For example, you may
use the History API.” (bản dịch) «Nếu đó links on AJAX các trang dùng đó character, thay đổi đó addresses để URLs không có này character. Ví dụ, bạn có thể dùng đó History API.»
Điều gì 2023 nguồn-code leak được tiết lộ về crawler
Trong January 2023, Yandex internal mã nguồn leaked — một well-corroborated event covered trên Search Engine Land, Search Engine Journal, và others. Treat này as leaked internal tài liệu, không một Yandex statement, nhưng điều này exposed real detail về cách đó crawler hoạt động. Theo Mike King SEL analysis: “Yandex’s documentation discusses a dual-distributed crawler system. One for real-time crawling called the ‘Orange Crawler’ and another for general crawling.” (bản dịch) «Yandex tài liệu discusses một dual-distributed crawler hệ thống. Một cho real-time crawling called đó ‘Orange Crawler’ và một sản phẩm khác cho chung crawling.» He drew một parallel để Google, mà “is said to have had an index stratified into three buckets, one for housing real-time crawl, one for regularly crawled and one for rarely crawled.” (bản dịch) «là đã nói để có đã có an chỉ mục stratified vào three buckets, một cho housing real-time crawl, một cho regularly được crawl và một cho rarely được crawl.» Cả hai engines, nói cách khác, xuất hiện để embrace segmented crawling driven by cách thường nội dung cập nhật.
Đó leak cũng tied crawling trực tiếp để kiến trúc trang web. Theo Dan Taylor SEJ coverage, “URLs that are reachable from the homepage have a ‘higher’ level of importance.” (bản dịch) «URLs đó là reachable từ đó homepage có một ‘cao hơn’ cấp độ of importance.» đó là một sạch bridge từ “crawler mechanics” (bản dịch) «crawler mechanics» để “why internal linking matters” (bản dịch) «vì sao liên kết nội bộ matters» — đó giống nhau crawl depth logic đó áp dụng để Googlebot.
(Cho quy mô: coverage noted đó widely-cited “1,922 ranking factors” (bản dịch) «1 922 xếp hạng factors» hình đã là cụ thể để một archive file, với đó fuller codebase reportedly containing far hơn trên multiple files — giữ một date và một nguồn on bất kỳ cụ thể number nếu bạn cite một.)
Vì sao YandexBot vẫn matters — và Cách phổ biến nó là trong robots.txt
Yandex global share là tiny, nhưng của nó Russia share không phải: ~71% trong Russia so với. Google ~27% as của June 2026 (StatCounter). đó stability là đểàn bộ reason YandexBot deserves tách biệt treatment. nếu bạn có bất kỳ Russia/CIS-facing business, blocking YandexBot forecloses dominant engine trong đó market.
Cách thường làm các trang ngay cả bother configuring cho điều này? Rarely, nhưng rising. Trong đó Web Almanac 2022 SEO chapter (I đã là một reviewer đó năm; I đã là lead tác giả of đó 2021 chapter): YandexBot appeared trong “just 0.5% of robots.txt files in 2021. By 2022, there was a six-fold increase, with 3% of files specifying Yandexbot.” (bản dịch) «chỉ 0,5% of robots.txt files trong 2021. By 2022, ở đó đã là một six-fold increase, với 3% of files specifying Yandexbot.» Nhỏ, nhưng một clear upward trend — và một hữu ích “how common is this in the wild” (bản dịch) «cách phổ biến là này trong đó wild» baseline.
nếu bạn’re deciding liệu để block nó, honest cách diễn đạt là business câu hỏi, không kỹ thuật một — và nó thực debate đó plays out trong quản trị viên web forums. cho rộng hơn mechanics YandexBot lives bên trong — phát hiện URL, crawl scheduler, kết xuất, và crawl-so với-chỉ mục-so với-xếp hạng distinctions — see crawling hub. và cho configuring Yandex cụ thể as part của Russia/CIS strategy, international-SEO và market-cụ thể-SEO material ties nó together.
AI summary
condensed take on Nâng cao version:
- YandexBot = Yandex Tìm kiếm main lập chỉ mục crawler — đó Yandex tương đương of
Googlebot/Bingbot. UA string:
Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/81.0.4044.268(không match on đó Chrome version — điều này thay đổi). - Hai robots.txt tokens, khác nhau scopes:
YandexBot= main lập chỉ mục bot chỉ;Yandex= đó rộng hơn bot family. MỘTUser-agent: Yandexblock làm Yandex bỏ quaUser-agent: *. Some Yandex bots có thể bỏ qua robots.txt hoàn toàn. - Sạch-param là Yandex-chỉ — consolidates URL parameters đó không thay đổi
nội dung (không Google/Bing tương đương). Điều này không require
Disallow, và đây là intersectional (có thể go anywhere trong đó file). - Crawl-delay là dead — Yandex đã dừng honoring điều này on February 22, 2018; dùng đó Tốc độ crawl tool thay vì. Some SEO các hướng dẫn vẫn wrongly claim Yandex hỗ trợ điều này. (Bing làm; Yandex không.)
Disallow≠noindex— disallowed các trang “can participate in Yandex search.” (bản dịch) «có thể participate trong Yandex tìm kiếm.» Dùngnoindex(meta hoặc HTTP header) để thực ra xóa một trang.- JavaScript kết xuất là beta và “at the bot’s discretion.” (bản dịch) «tại đó bot discretion.» Đó 2023 leak suggested “no separate rendering system for JavaScript” (bản dịch) «không tách biệt kết xuất hệ thống cho JavaScript» — so đây là selective và simpler hơn Google hai-wave pipeline. Ưu tiên SSR/pre-kết xuất; tránh hash-bang URLs (dùng đó History API).
- 2023 nguồn-code leak (leaked, không chính thức): được tiết lộ một dual-crawler hệ thống (real-time “Orange Crawler” (bản dịch) «Orange Crawler» + chung crawler) và đó homepage-reachable URLs carry cao hơn importance (crawl-depth as một tín hiệu).
- Verify với reverse-thì-forward DNS để một
yandex.ru/yandex.net/yandex.comhost — đó giống nhau FCrDNS technique Google và Bing dùng — không bao giờ đó spoofable người dùng-agent. - Vì sao care: ~71% Russia tìm kiếm share (so với. Google ~27%, June 2026). Tiny globally, dominant regionally. Trong robots.txt: ~0,5% of files (2021) → 3% (2022) theo đó Web Almanac.
Tài liệu chính thức
Chính-nguồn tài liệu, mostly từ Yandex Quản trị viên web Help, với Google/Bing contrast links. Heads lên: several Yandex Quản trị viên web tool các trang (Tốc độ crawl, robots.txt trình phân tích) là app các trang đó cần logged-trong trình duyệt để see fully, và Google/Bing verification các trang là JavaScript-được kết xuất — xác nhận details trực tiếp nếu link bounces.
Yandex
- Người dùng-agent directive — đó
YandexBotso với.Yandextoken scoping và precedence rules. - Cách kiểm tra rằng một robot belongs để Yandex — đó đầy đủ người dùng-agent string và đó reverse-DNS verification phương thức.
- Dùng robots.txt — đó directives Yandex recognizes, đó 500 KB / HTTP 200 file requirements, và đó
Disallow≠ chỉ mục note. - Đó Sạch-param directive — Yandex unique parameter-consolidation directive (syntax + worked ví dụ).
- Đó Crawl-delay directive — đó trang stating Crawl-delay đã được đã bỏ qua since February 22, 2018.
- Site tốc độ crawl — đó Crawl-delay replacement trong Yandex Quản trị viên web.
- Lập chỉ mục các trang với JavaScript (β) — đó “at the bot’s discretion” (bản dịch) «tại đó bot discretion» default và đó SSR khuyến nghị.
- Lập chỉ mục AJAX các trang — JS execution on original URLs và đó History-API khuyến nghị.
Google / Bing (cho contrast)
- Verify Các yêu cầu từ Google Các crawler và Fetchers — Google reverse/forward-DNS phương thức (so với
googlebot.com/google.com/googleusercontent.com) — giống nhau FCrDNS pattern Yandex dùng. - Mà các crawler làm Bing sử dụng? — Bing crawler list; Bing verifies so với
*.search.msn.comvà vẫn honorscrawl-delay(Yandex không).
Quotes từ nguồn
On—record statements từ Yandex own tài liệu, plus leaked-internal material relayed qua ngành coverage. mỗi link là deep link đó jumps để quoted passage. Yandex tài liệu là English translations published by Yandex itself — quoted verbatim từ những điều đó các trang.
Yandex — robots.txt tokens và precedence
- “If the
User-agent: Yandexstring is detected, theUser-agent: *string is ignored.” (bản dịch) «Nếu đó string là detected, đó string là đã bỏ qua.» Nhảy đến trích dẫn - “Some Yandex robots may ignore directives in
robots.txt, including those forUser-agent: Yandex.” (bản dịch) «Some Yandex robots có thể bỏ qua directives trong , including những cho .» Nhảy đến trích dẫn - “
YandexAdditionalBot… Helps process robots.txt to prevent page content from appearing in Search with Yandex AI responses. Applies to the pages that have been indexed by the primary crawler.” (bản dịch) «… Helps xử lý robots.txt để ngăn trang nội dung từ appearing trong Tìm kiếm với Yandex AI các phản hồi. Áp dụng để đó các trang đó đã được lập chỉ mục by đó chính crawler.» Nhảy đến trích dẫn
Yandex — verifying bot
- “Some robots can disguise themselves as Yandex robots by indicating the relevant User Agent. You can check the authenticity of a robot using a reverse DNS lookup.” (bản dịch) «Some robots có thể disguise themselves as Yandex robots by indicating đó relevant Người dùng Agent. Bạn có thể kiểm tra đó authenticity of một robot dùng một reverse DNS lookup.» Nhảy đến trích dẫn
- “Check whether the host belongs to Yandex. All Yandex robots have names ending in
yandex.ru,yandex.netoryandex.com.” (bản dịch) «Kiểm tra xem đó host belongs để Yandex. All Yandex robots có names ending trong , hoặc .» Nhảy đến trích dẫn - “If the IP addresses do not match, it means that the host name is fake.” (bản dịch) «Nếu đó IP addresses không match, điều này có nghĩa là đó host name là fake.» Nhảy đến trích dẫn
Yandex — robots.txt và Sạch-param
- “Pages restricted in
robots.txtcan participate in Yandex search. To remove pages from search, specify thenoindexdirective in the HTML code of the page or configure the HTTP header.” (bản dịch) «Các trang restricted trong có thể participate trong Yandex tìm kiếm. Để xóa các trang từ tìm kiếm, specify đó directive trong đó HTML code of đó trang hoặc configure đó HTTP header.» Nhảy đến trích dẫn - “The Yandex robot uses this directive to avoid reloading duplicate information. This improves the robot’s efficiently and reduces the server load.” (bản dịch) «Đó Yandex robot dùng này directive để tránh reloading duplicate information. Này improves đó robot efficiently và reduces đó máy chủ load.» [sic — “efficiently” là Yandex own typo] Nhảy đến trích dẫn
- “The Clean-param directive does not require mandatory combination with the Disallow directive.” (bản dịch) «Đó Sạch-param directive không yêu cầu mandatory combination với đó Disallow directive.» Nhảy đến trích dẫn
- “Do not restrict such pages in
robots.txt, or the Yandex bot can’t index them and detect your instructions.” (bản dịch) «Không restrict such các trang trong , hoặc đó Yandex bot không thể chỉ mục them và detect của bạn instructions.» Nhảy đến trích dẫn
Yandex — Crawl-delay và JavaScript
- “From February 22, 2018, Yandex doesn’t take into account the Crawl-delay directive.” (bản dịch) «Từ February 22, 2018, Yandex không take vào account đó Crawl-delay directive.» Nhảy đến trích dẫn
- “Prohibit rendering if SSR (Server-Side Rendering) or pre-rendering is implemented on the site.” (bản dịch) «Prohibit kết xuất nếu SSR (Máy chủ-Side Kết xuất) hoặc pre-kết xuất là implemented on đó site.» Nhảy đến trích dẫn
- “When indexing an AJAX site, the Yandex bot scans the original URLs and executes JavaScript code on them.” (bản dịch) «Khi lập chỉ mục an AJAX site, đó Yandex bot scans đó original URLs và executes JavaScript code on them.» Nhảy đến trích dẫn
** 2023 Yandex nguồn-code leak** (leaked internal tài liệu, relayed qua ngành coverage — không chính thức Yandex statement)
- “Yandex has no separate rendering system for JavaScript. They say this in their documentation and, although they have Webdriver-based system for visual regression testing called Gemini, they limit themselves to text-based crawl.” (bản dịch) «Yandex có không tách biệt kết xuất hệ thống cho JavaScript. They chẳng hạn này trong của họ tài liệu và, although they có Webdriver-based hệ thống cho visual regression kiểm thử called Gemini, they limit themselves để text-based crawl.» — Mike King, iPullRank, qua Search Engine Land. Đọc bài đưa tin
- “Yandex’s documentation discusses a dual-distributed crawler system. One for real-time crawling called the ‘Orange Crawler’ and another for general crawling.” (bản dịch) «Yandex tài liệu discusses một dual-distributed crawler hệ thống. Một cho real-time crawling called đó ‘Orange Crawler’ và một sản phẩm khác cho chung crawling.» — Search Engine Land. Đọc bài đưa tin
- “There’s nothing new to suggest Yandex can crawl JavaScript yet outside of already publicly documented processes.” (bản dịch) «có không có gì new để suggest Yandex có thể crawl JavaScript tuy vậy bên ngoài of đã publicly được ghi lại xử lý.» — Dan Taylor, qua Search Engine Journal. Đọc bài đưa tin
Web Almanac 2022, SEO chapter (I là reviewer)
- “Yandexbot was specified in just 0.5% of robots.txt files in 2021. By 2022, there was a six-fold increase, with 3% of files specifying Yandexbot.” (bản dịch) «Yandexbot đã là specified trong chỉ 0,5% of robots.txt files trong 2021. By 2022, ở đó đã là một six-fold increase, với 3% of files specifying Yandexbot.» Nhảy đến trích dẫn
nên I block, cho phép, hoặc throttle YandexBot?
YandexBot câu hỏi là gần như luôn business câu hỏi wearing kỹ thuật costume. Hoạt động qua nó.
What to do about YandexBot in your logs
YandexBot myths và mistakes để tránh
traps đó come lên phần lớn thường — several là widely-repeated và worth correcting:
- “Crawl-delay slows YandexBot down.” (bản dịch) «Crawl-delay làm chậm YandexBot xuống.» Không. Yandex đã dừng honoring
Crawl-delayon February 22, 2018 (“Yandex doesn’t take into account the Crawl-delay directive” (bản dịch) «Yandex không take vào account đó Crawl-delay directive»). Some SEO các hướng dẫn vẫn claim Yandex hỗ trợ điều này — đó là stale. Dùng đó Tốc độ crawl setting trong Yandex Quản trị viên web thay vì. (Làm-thay vì: verify crawl-behavior claims so với Yandex trực tiếp tài liệu, và throttle qua Tốc độ crawl, không Crawl-delay.) - “
Disallowin robots.txt removes the page from Yandex.” (bản dịch) «trong robots.txt xóa đó trang từ Yandex.» Không — Yandex own tài liệu chẳng hạn disallowed các trang “can participate in Yandex search.” (bản dịch) «có thể participate trong Yandex tìm kiếm.»Disallowdừng crawling, không lập chỉ mục. (Làm-thay vì: dùngnoindextrong đó HTML hoặc an HTTP header để thực ra exclude một trang; let Yandex crawl điều này so điều này có thể see đó tag.) - “
User-agent: YandexBotandUser-agent: Yandexare the same thing.” (bản dịch) «và là cùng một điều.» Không —YandexBottargets chỉ đó main lập chỉ mục bot;Yandextargets đó rộng hơn family. Và mộtUser-agent: Yandexblock làm Yandex bỏ qua của bạnUser-agent: *block hoàn toàn. (Làm-thay vì: pick đó token đó matches đó phạm vi bạn thực ra intend, và không assume*rules reach Yandex bots.) - “YandexBot renders JavaScript just like Googlebot.” (bản dịch) «YandexBot renders JavaScript chỉ như Googlebot.» Không established. Yandex kết xuất là beta và “at the bot’s discretion,” (bản dịch) «tại đó bot discretion,» và đó 2023 leak suggested “no separate rendering system for JavaScript.” (bản dịch) «không tách biệt kết xuất hệ thống cho JavaScript.» (Làm-thay vì: serve JS-phụ thuộc nội dung qua SSR/pre-kết xuất thay vì assuming client-side kết xuất sẽ là được lập chỉ mục.)
- “The user-agent string proves it’s YandexBot.” (bản dịch) «Người dùng-agent string proves đây là YandexBot.» Không — đây là trivially spoofed.
(Làm-thay vì: verify với reverse-thì-forward DNS để một
yandex.ru/yandex.net/yandex.comhost, chính xác as bạn’d verify Googlebot hoặc Bingbot.) - “Any YandexBot traffic on a non-Russian site is inherently suspicious/fake.” (bản dịch) «Bất kỳ YandexBot traffic on một non-Russian site là inherently suspicious/fake.»
Không nhất thiết — legitimate YandexBot crawl globally-facing các trang cũng, không chỉ
.rudomains. (Làm-thay vì: verify by DNS trước deciding đây là spoofed; một real Yandex host là real regardless of của bạn audience.) - “The 2023 leak proved Yandex has a secret superior JS crawler.” (bản dịch) «Đó 2023 leak proved Yandex có một secret superior JS crawler.» Không — coverage concluded có “nothing new to suggest Yandex can crawl JavaScript yet outside of already publicly documented processes.” (bản dịch) «không có gì new để suggest Yandex có thể crawl JavaScript tuy vậy bên ngoài of đã publicly được ghi lại xử lý.» (Làm-thay vì: không over-đọc đó leak; điều này corroborated đó limited-kết xuất picture, điều này đã không overturn điều này.)
YandexBot — bảng tra nhanh
Người dùng-agent string (main lập chỉ mục bot)
Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/81.0.4044.268Match on YandexBot token, không Chrome/81.0.4044.268 version — Yandex
nói trình duyệt version có thể thay đổi.
** hai robots.txt tokens**
| Token | Phạm vi |
|---|---|
User-agent: YandexBot | main lập chỉ mục bot chỉ |
User-agent: Yandex | Yandex rộng hơn bot family |
User-agent: Yandex block gây ra Yandex để bỏ qua của bạn User-agent: * block.
Some Yandex bots có thể bỏ qua robots.txt hoàn toàn.
Directives Yandex recognizes
| Directive | Điều gì nó làm |
|---|---|
User-agent | Mà robot rules apply để |
Disallow | Prohibits crawling (không lập chỉ mục) |
Allow | Cho phép crawling/lập chỉ mục của sections hoặc các trang |
Sitemap | Path để sitemap file |
Clean-param | Yandex-chỉ — bỏ qua listed URL parameters (không Google/Bing tương đương) |
Sạch-param nhanh form
User-agent: Yandex
Clean-param: ref /some_dir/get_book.plConsolidates ?ref=… variants để một URL. không cần Disallow; intersectional
(có thể go anywhere trong file).
Verify so với. Google so với. Bing (all sử dụng forward-confirmed reverse DNS)
| Engine | Reverse-DNS host phải end trong |
|---|---|
| Yandex | yandex.ru / yandex.net / yandex.com |
googlebot.com / google.com / googleusercontent.com | |
| Bing | *.search.msn.com |
Fast facts
- Crawl-delay: dead since Feb 22, 2018 — dùng đó Tốc độ crawl tool. (Bing
vẫn honors
crawl-delay; Yandex không.) robots.txtphải được namedrobots.txt, là ≤ 500 KB, và trả về HTTP 200.- JavaScript kết xuất: beta, “at the bot’s discretion” (bản dịch) «tại đó bot discretion» — ưu tiên SSR.
- Russia share: ~71% (so với. Google ~27%), June 2026 (StatCounter).
- Trong robots.txt: 0,5% (2021) → 3% (2022) of files (Web Almanac).
Verify bot là thực sự YandexBot
người dùng-agent là trivially spoofed, so xác nhận với reverse-DNS lookup (phải end
trong Yandex domain) followed by forward-DNS lookup (phải resolve lại để giống nhau
IP). Đây là chính xác FCrDNS pattern bạn’d sử dụng cho Googlebot hoặc Bingbot — chỉ
domain suffixes differ (yandex.ru / yandex.net / yandex.com).
macOS / Linux
# 1) Reverse-DNS the IP from your logs — the host must end in
# yandex.ru, yandex.net, or yandex.com
host 5.255.253.1
# → 1.253.255.5.in-addr.arpa domain name pointer <something>.yandex.com (illustrative)
# 2) Forward-DNS that hostname back — it must resolve to the same IP
host <the-hostname-from-step-1>.yandex.comWindows
nslookup 5.255.253.1
nslookup <the-hostname-from-step-1>.yandex.comnếu reverse lookup không end trong Yandex domain, hoặc forward lookup không trả về gốc IP, nó không phải YandexBot — drop hoặc rate-limit nó.
Extract verified-so với-spoofed YandexBot hits từ log file
nhanh shell truyền để pull “YandexBot” lines và kiểm tra mỗi nguồn IP PTR. Adjust trường positions cho của bạn log format (điều này assumes phổ biến combined format với IP đầu tiên).
# Pull unique IPs that claimed to be YandexBot, then reverse-resolve each
grep -i 'YandexBot' access.log \
| awk '{print $1}' | sort -u \
| while read ip; do
host="$(host "$ip" 2>/dev/null | awk '/pointer/{print $NF}' | sed 's/\.$//')"
case "$host" in
*.yandex.ru|*.yandex.net|*.yandex.com) echo "REAL $ip $host" ;;
"" ) echo "NO-PTR $ip" ;;
* ) echo "FAKE $ip $host" ;;
esac
doneREAL lines vẫn deserve forward-lookup xác nhận cho bất cứ điều gì bạn’ll act on
(host "$host" nên trả về gốc IP), nhưng điều này triages obvious
impostors đầu tiên.
Regex để match YandexBot token trong người dùng-agent
Match sản phẩm token, không Chrome version (mà thay đổi). Case-insensitive:
YandexBot/\d+(\.\d+)?Rộng hơn “any Yandex robot” (bản dịch) «bất kỳ Yandex robot» match (catches YandexImages, YandexMobileBot, etc.):
Yandex[A-Za-z]*/\dRemember: matching người dùng-agent là cần thiết nhưng chư đủ — pair bất kỳ regex match với DNS kiểm tra trên trước khi trusting nó.
robots.txt starting point cho Yandex
# Slow/duplicate-parameter cleanup for all Yandex bots
User-agent: Yandex
Clean-param: utm_source&utm_medium&utm_campaign /
# Block only the main indexing bot from a low-value space
User-agent: YandexBot
Disallow: /internal-search/
Sitemap: https://example.com/sitemap.xmlRemember: Disallow chặn crawling, không lập chỉ mục — sử dụng noindex để xóa
trang từ Yandex tìm kiếm. để fully block all Yandex bots (some bỏ qua robots.txt),
block by verified IP tại máy chủ.
YandexBot readiness checklist
nhanh truyền để xác nhận bạn’re xử lý YandexBot có chủ ý, không by accident:
- bạn đã decided liệu Yandex matters cho trang web của bạn (bất kỳ Russia/CIS audience hoặc business?) — đó toàn bộ block/cho phép câu hỏi hinges on này.
- Bot verification dùng reverse + forward DNS để một
yandex.ru/yandex.net/yandex.comhost — không đó spoofable người dùng-agent, và không một hardcoded IP list. - bạn là dùng đó right robots.txt token cho của bạn intent:
YandexBot(main lập chỉ mục bot chỉ) so với.Yandex(đó rộng hơn family). - Bạn know một
User-agent: Yandexblock làm Yandex bỏ qua của bạnUser-agent: *block. - bạn là không relying on
Crawl-delay(dead since Feb 22, 2018) — throttle qua đó Tốc độ crawl tool trong Yandex Quản trị viên web thay vì. - bạn là không dùng
Disallowđể deindex — đó lànoindex’s job (với crawling được phép). - Duplicate URL parameters là consolidated với Sạch-param nơi relevant
(điều này không cần
Disallow, và điều này có thể go anywhere trong đó file). - JS-phụ thuộc nội dung là phân phối qua SSR/pre-kết xuất, since Yandex kết xuất là beta và “at the bot’s discretion.” (bản dịch) «tại đó bot discretion.»
- AJAX routes dùng sạch URLs / đó History API, không hash-bang (
#!) URLs. - Cho đầy đủ exclusion of all Yandex bots (some bỏ qua robots.txt), bạn là blocking by verified IP tại đó máy chủ, không chỉ trong robots.txt.
Tools cho verifying và monitoring YandexBot
hai jobs đó come lên phần lớn với YandexBot là confirming hit là thực và seeing Cách nhiều nó thực ra crawling bạn. Bắt đầu ở đây:
- Googlebot Verifier — paste an IP từ của bạn logs và điều này chạy
đó forward-confirmed reverse-DNS kiểm tra cho bạn (đó giống nhau phương thức Yandex documents
cho confirming một
yandex.ru/yandex.net/yandex.comhost), naming đó real network owner khi đây là một spoofer thay vì. Nhanh hơn đang chạyhost/nslookupby hand cho mỗi suspicious hit, và điều này cũng covers Googlebot, Bingbot, và đó major AI các crawler nếu bạn là kiểm tra một mixed log. - Log File Trình phân tích — drop trong một máy chủ access log (nginx, Apache, IIS/W3C, hoặc JSON) và see YandexBot thực tế crawl footprint: cách nhiều of trang web của bạn đây là hitting, mà sections, status-code waste, và một spoofer báo cáo flagging IPs đó claim để là YandexBot nhưng không kiểm tra out. Hữu ích cho đó “is this crawling actually straining my server” (bản dịch) «là này crawling thực ra straining my máy chủ» câu hỏi từ đó decision tree trên — mọi thứ chạy trong trình duyệt của bạn, không có gì là uploaded.
từ Yandex itself
- Yandex Quản trị viên web — Tốc độ crawl setting ( Crawl-delay replacement) và robots.txt trình phân tích trực tiếp ở đây; bạn’ll cần verified, logged-trong Yandex Quản trị viên web account để sử dụng either.
các tài nguyên worth của bạn time
My related writing
- được lập chỉ mục, though Bị chặn bởi robots.txt — Vì sao robots-blocked URL vẫn nhận được lập chỉ mục; giống nhau
Disallow≠noindextrap áp dụng để Yandex. - Robots.txt và SEO: Mọi thứ bạn cần Know — chung robots.txt reference (note: của nó Yandex
Crawl-delayclaim là outdated, theo Yandex own dated tài liệu — good ví dụ của verifying so với nguồn). - Story của Blocking 2 Cao-Xếp hạng các trang với Robots.txt — my đầu tiên-party thử nghiệm on Điều gì thực ra happens Khi bạn block xếp hạng các trang.
- Đáp ứng New Web Các crawler: AI Bots là Closing trong on công cụ tìm kiếm Bots — Cách crawler cast trong của bạn nhật ký có changed.
- SEO Bots đó ~140 Million Websites Block phần lớn — my nghiên cứu với Xibeijia Guan on robots.txt block rates (nó covers Western SEO-tool bots, không YandexBot — cited cho methodology và as contrast để Cách rarely các trang configure cho Yandex).
My speaking
- Cách Tìm kiếm Hoạt động (SlideShare) — my walkthrough of crawling, kết xuất, lập chỉ mục, và xếp hạng. (Standing disclaimer áp dụng: “This is my understanding of systems… not going to be 100% complete or accurate.” (bản dịch) «Này là my understanding of các hệ thống… không going để là 100% hoàn tất hoặc chính xác.»)
Từ khoảng đó ngành
- Yandex scrapes Google và other SEO learnings từ đó mã nguồn leak (Mike King / iPullRank, Search Engine Land, Jan 30, 2023) — đó “Orange Crawler” (bản dịch) «Orange Crawler» dual-crawler hệ thống và “no separate rendering system for JavaScript” (bản dịch) «không tách biệt kết xuất hệ thống cho JavaScript» findings.
- Yandex Dữ liệu Leak: Đó Xếp hạng Factors & Đó Myths We Được tìm thấy (Dan Taylor, Search Engine Journal, Feb 1, 2023) — đó crawl-depth-as-importance tín hiệu và đó “nothing new on JavaScript crawling” (bản dịch) «không có gì new on JavaScript crawling» conclusion.
- Yandex ‘leak’ reveals 1 922 tìm kiếm xếp hạng factors (Search Engine Land) — context và dating cho đó leak quy mô.
- Làm đó Yandex Code Leak Tell Us Bất cứ điều gì Về Google? (seoClarity) — một measured cross-engine đọc of đó leak.
- Đó Ultimate Hướng dẫn để Yandex SEO (Search Engine Journal) — rộng hơn Yandex-optimization context beyond đó crawler.
- Web Almanac 2022 — SEO chapter (HTTP Archive) — đó nguồn of đó YandexBot-trong-robots.txt adoption stat.
Số liệu worth citing
- ~71% Russia tìm kiếm share — Yandex share of đó Russian tìm kiếm market versus Google ~27%, theo StatCounter (dữ liệu reported cho June 2026; nhạy cảm với thời gian, so cite đó month/năm và expect drift). Này là đó entire reason YandexBot warrants tách biệt treatment từ “block all non-Google bots.” (bản dịch) «block all non-Google bots.» Nguồn
- 0,5% → 3% of robots.txt files — YandexBot went từ đang specified trong “just 0.5% of robots.txt files in 2021” (bản dịch) «chỉ 0,5% of robots.txt files trong 2021» để “3% of files specifying Yandexbot” (bản dịch) «3% of files specifying Yandexbot» trong 2022 — một six-fold increase, though vẫn nhỏ (Web Almanac 2022 SEO chapter, mà I reviewed). Hữu ích as một lịch sử baseline cho “how common is this in the wild.” (bản dịch) «cách phổ biến là này trong đó wild.» Nguồn
- February 22, 2018 — đó date Yandex đã dừng honoring
Crawl-delay, theo của nó own tài liệu. MỘT precise, citable deprecation date để counter stale các hướng dẫn claiming Yandex vẫn hỗ trợ đó directive. Nguồn
Tự kiểm tra: YandexBot
Five nhanh các câu hỏi on YandexBot, của nó robots.txt behavior, và Cách nó compares để Googlebot và Bingbot. Pick câu trả lời cho mỗi, sau đó kiểm tra.
Nhật ký thay đổi
Đã cập nhật 8 thg 8, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.
Đã cập nhật 18 thg 7, 2026.
Tóm tắt biên tập và chi tiết thay đổi đã ghi nhận.Chi tiết thay đổi
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
-
Ghi chú thay đổi chi tiết hiện chỉ có bằng tiếng Anh.
Không thể so sánh đầy đủ — không có bản lưu trước đó cho lần sửa đổi này.