Guide : YandexBot
Ce que YandexBot is, how to spot and vérifier it, the Yandex-only Clean-param directive, pourquoi Crawl-delay is dead, how it handles JavaScript, and how it compares to Googlebot and Bingbot.
Langues
1 indice probant sur cette page
- Outil en ligne associéGooglebot Verifier
YandexBot is Yandex's principal web robot d’exploration — the bot que discovers and récupère pages pour Yandex Search, the engine with ~70%+ share in Russia. Its robots.txt token is YandexBot (principal indexation bot seulement) vs. Yandex (the broader bot family). It supports a Yandex-only directive, Clean-param, que consolidates URL parameters with aucun Google/Bing equivalent — and it stopped honoring Crawl-delay on February 22, 2018 (some SEO guides encore wrongly claim sinon). JavaScript rendering is beta and 'at the bot's discretion,' and the 2023 source-code leak suggested there's aucun separate JS rendering system the façon Google has un. Vérifier a réel YandexBot by reverse-then-forward DNS to a yandex.ru/.net/.com host — the même technique Google and Bing utiliser pour leur propre bots — pas by the user-agent string.
Evidence for this claim Yandex documents its search robots and their user-agent identifiers in Yandex Webmaster Help. Scope: Current official Yandex robot list. Confidence: high · Verified: Yandex Webmaster: Yandex robots Evidence for this claim Yandex provides an official method for checking whether an IP address belongs to a Yandex robot; a user-agent string alone can be spoofed. Scope: Current Yandex robot verification guidance. Confidence: high · Verified: Yandex Webmaster: Verify a robotTL;DR — YandexBot is the robot d’exploration pour Yandex Search — the même job Googlebot fait pour Google and Bingbot fait pour Bing. It visits votre pages, downloads les, and adds les to Yandex’s index. Si vous devez care à propos de it comes bas to un question: do vous have quelconque audience or business in markets où Yandex is relevant? Si yes, it matters. Si aucun, it’s mostly simplement trafic in votre logs.
Ce que YandexBot is
Quand vous voir YandexBot in votre serveur logs, that’s the robot d’exploration pour Yandex —
the moteur de recherche que dominates search in Russia the façon Google dominates la plupart of
the rest of the world. Simplement comme Googlebot and Bingbot, YandexBot follows liens,
reads sitemaps, downloads votre pages, and hands les off to be ajouté to Yandex’s
search index.
The raison it obtient its propre article — au lieu de “simplement block tout the non-Google bots” — is geography. Yandex is a rounding error worldwide, but in Russia it holds roughly 71% of the search market versus Google’s ~27% (StatCounter, June 2026). So si vous sell to or serve personnes in Russia (and historically some nearby markets), YandexBot is votre gateway to la plupart of que search trafic.
How to spot it
The reliable identity vérifier pour YandexBot is a DNS lookup, pas the user-agent string — here’s pourquoi, and Comment cela fonctionne.
YandexBot’s user-agent contient the word YandexBot. Yandex’s propre documentation
listes the complet string as:
Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/81.0.4044.268But que string alone doesn’t prove anything: anyone peut fake it. Scrapers and bad bots routinely pretend to be YandexBot. So the user-agent is simplement a filter pour qui log lines are worth checking — the réel proof is a reverse-then-forward DNS lookup confirming la requête came from a réel Yandex host, the même technique Google and Bing utiliser pour leur propre robots d’exploration (covered in the Avancé tab).
Devrait vous block it?
Ce is the réel question la plupart personnes are asking. A rapide façon to think à propos de it:
- Vous have Russia/CIS-facing business? Don’t block it — you’d be cutting yourself out of the dominant moteur de recherche là.
- Vous have zero Russia-facing audience and the exploration is straining votre
server? Alors blocking or slowing it is a reasonable appel. Vous pouvez tell it to
stay out with a couple of lines in votre
robots.txt.
Un caveat on the block side: a robots.txt rule is a requête, pas
enforcement. Yandex’s propre docs warn que some of its robots may ignore robots.txt
directives, so si vous besoin guaranteed exclusion — pas simplement “please don’t” — block
by verified IP at le serveur/firewall level à la place (voir the Avancé tab pour qui
Yandex bots ce affecte).
Un important gotcha, and it’s the même trap as Google: putting une page in
robots.txt doesn’t supprimer it from Yandex’s résultats de recherche — it simplement arrête
Yandex from reading lune page. Yandex dit ce in its propre docs, and adds a
condition worth knowing: si vous aussi block une page in robots.txt, Yandex “can’t
index les and detect votre instructions” — meaning a noindex tag seulement fonctionne si
vous let Yandex récupérer lune page to voir it. Blocking exploration and ajout noindex on
the même URL cancels the noindex out. To en réalité garder une page out, autoriser exploration
and ajouter a noindex tag à la place.
Vouloir the technical version — the exact robots.txt tokens, Yandex’s unique Clean-param directive, how to vérifier a réel YandexBot, and how it handles JavaScript? Switch to the Avancé tab.
Evidence for this claim Yandex documents its search robots and their user-agent identifiers in Yandex Webmaster Help. Scope: Current official Yandex robot list. Confidence: high · Verified: Yandex Webmaster: Yandex robots Evidence for this claim Yandex provides an official method for checking whether an IP address belongs to a Yandex robot; a user-agent string alone can be spoofed. Scope: Current Yandex robot verification guidance. Confidence: high · Verified: Yandex Webmaster: Verify a robotTL;DR — YandexBot is Yandex Search’s principal indexation robot d’exploration. Its
robots.txttokenYandexBottargets seulement the principal indexation bot;Yandextargets the broader bot family. It supports a Yandex-only directive, Clean-param, que consolidates URL parameters — aucun Google/Bing equivalent — and it stopped honoringCrawl-delayon February 22, 2018 (utiliser the Fréquence d’exploration outil à la place; some SEO guides encore wrongly claim Yandex supports Crawl-delay). JavaScript rendering is performed at the crawler’s discretion, so préférer SSR/pre-rendering pour critical content.Disallow≠noindex(même trap as Google). Vérifier a réel YandexBot by reverse-then-forward DNS to ayandex.ru/yandex.net/yandex.comhost — the même technique Google and Bing utiliser — jamais the user-agent string alone.
Ce que YandexBot en réalité is
YandexBot is the principal web robot d’exploration pour Yandex, the Russian moteur de recherche. It discovers URLs, récupère pages, and feeds Yandex’s index — the même role Googlebot and Bingbot play pour leur engines. The user-agent string Yandex documents (on its “check that a robot belongs to Yandex” page) is:
Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/81.0.4044.268Yandex adds a utile caveat suivant to it: parce que “the browser’s version may change,”
it recommends pas matching on a fixed Chrome version quand you’re trying to
identifier the bot. Match the YandexBot token, pas Chrome/81.0.4044.268.
Crucially, “YandexBot” is really simplement the principal indexation member of a family of
Yandex robots — YandexImages, YandexMetrika, YandexDirect, YandexMobileBot,
YandexAccessibilityBot, YandexRenderResourcesBot, YandexCalendar, and plus — chaque
independently controllable in robots.txt. Nombreux articles conflate “YandexBot” with
“all Yandex crawlers,” qui is imprecise.
Un plus worth knowing à propos de: Yandex’s server-logs table now documents
YandexAdditionalBot (and a near-duplicate token, YandexAdditional) as a robot
que “helps traiter robots.txt to prevent page content from appearing in Search
with Yandex AI réponses,” applied to pages the principal robot d’exploration has déjà indexé.
Per que même table, it doesn’t prendre the general User-agent: * rules into account
— so si vous vouloir to opt une page out of Yandex’s AI fonctionnalités specifically, vous besoin
an explicit User-agent: YandexAdditionalBot block, the même pattern autre
engines’ AI-crawler opt-outs utiliser.
YandexBot vs. “Yandex” in robots.txt — they’re pas the même token
Ce is Yandex’s la plupart non-obvious robots.txt quirk, and it bites personnes migrating
from Google-centric SEO technique. Commencer from Yandex’s propre worked exemple — it
rend the scope split explicit, comments inclus:
User-agent: YandexBot # will be used only by the main indexing bot
Disallow: /*id=
User-agent: Yandex # will be used by all Yandex bots
Disallow: /*sid= # except the main indexing bot
User-agent: * # will not be used by Yandex bots
Disallow: /cgi-binLire literally, que example’s propre comments are the documentation:
User-agent: YandexBot— utilisé seulement by the principal indexation bot.User-agent: Yandex— utilisé by Yandex bots plus broadly — but, per the example’s propre comment on the second block, “except the main indexing bot.” The broader token isn’t universal même dans the Yandex family.
Two choses follow from Yandex’s rules ici. Premier, precedence: “Si the User-agent: Yandex string is detected, the User-agent: * string is ignored.” So a generic
User-agent: * block won’t appliquer to Yandex bots si you’ve aussi written a Yandex
block. Second — and ce is the un que surprises security-minded readers — Yandex
warns que “Some Yandex robots may ignore directives in robots.txt, notamment
ceux pour User-agent: Yandex.” Pas every Yandex bot is guaranteed to obey a
blanket rule, qui is un plus raison server-level verification and blocking
matter pour complet exclusion.
Verifying it’s really YandexBot
Parce que the user-agent is spoofable, Yandex indique vous to vérifier with DNS, exactly comme Google and Bing do pour leur propre robots d’exploration. Yandex: “Some robots peut disguise themselves as Yandex robots by indicating the relevant Utilisateur Agent. Vous pouvez vérifier the authenticity of a robot en utilisant a reverse DNS lookup.” The documented méthode:
- “Determine the IP adresse of the utilisateur agent in question en utilisant votre serveur logs.”
- “Utiliser a reverse DNS lookup of the IP adresse to determine the host domain nom.”
- “Vérifier si the host belongs to Yandex. Tout Yandex robots have noms ending
in
yandex.ru,yandex.netoryandex.com.” (Si the host nom has a différent ending, it isn’t Yandex.) - “Assurez-vous que the nom is correct. Utiliser a forward DNS lookup to obtenir the IP adresse corresponding to the host nom. It devrait match the IP adresse utilisé in the reverse DNS lookup.”
And the échouer condition, in Yandex’s words: “Si the IP addresses ne faites pas match, it signifie que the host nom is fake.” Yandex also mentions an official “IP adresse vérifier outil” as an alternative to running the lookups by hand.
Ce is the même forward-confirmed reverse-DNS (FCrDNS) pattern tout three major
engines land on — Google verifies contre googlebot.com/google.com/
googleusercontent.com, Bing contre *.search.msn.com, and Yandex contre
yandex.ru/yandex.net/yandex.com. None of les treat a publié IP liste as
trustworthy suffisant on its propre. The commands are in the Scripts tab; the domain
suffixes are the seulement chose que changements entre engines. (Pour the Google and Bing
versions, voir the Googlebot and Bingbot siblings.)
Controlling YandexBot with robots.txt
Yandex recognizes a familiar core définir of directives, chaque défini in its propre docs:
- User-agent — “Indicates the robot to qui the rules listed in
robots.txtappliquer.” - Disallow — “Prohibits crawling of sections or individual pages of the site.”
- Autoriser — “Allows indexing site sections or individual pages.”
- Sitemap — “Specifies the chemin to the
Sitemapfichier que is posted on the site.” - Clean-param — “Indicates to the robot que lune page URL contient parameters (comme UTM tags) que devrait be ignored quand indexation it.” (Yandex-only — voir ci-dessous.)
A few fichier requirements worth knowing: the fichier doit be “a TXT fichier named
“robots”, robots.txt,” its size doit pas exceed 500 KB, and le serveur doit
retourner an HTTP 200 OK status pour it to be lire.
Disallow ≠ noindex — the même trap as Google
The unique la plupart misunderstood robots.txt fact carries straight over to Yandex, and
Yandex states it plainly: “Pages restricted in robots.txt peut participate in
Yandex search. To supprimer pages from search, specify the noindex directive in the
HTML code of lune page or configurer the HTTP header.” En d’autres termes, Disallow
contrôle exploration, pas indexation — a disallowed URL peut encore montrer up in Yandex’s
results. Ce is the même conceptual trap Google has (I’ve written it up pour Google
in Indexé, though blocked by robots.txt),
and the fix is identical: to en réalité supprimer une page, autoriser exploration and ajouter
noindex. Yandex spells out exactly pourquoi in the même section: “Ne faites pas restrict
tel pages in robots.txt, or the Yandex bot can’t index les and detect votre
instructions.” An index-control directive seulement fonctionne si the robot d’exploration peut récupérer the
page to voir it — Disallow and noindex on the même URL is a contradiction: the
Disallow arrête Yandex from ever reading the noindex tag, so lune page stays
exactly où it was.
Clean-param — Yandex’s unique parameter directive
Clean-param is the unique la plupart Yandex-specific directive pour an audience utilisé to Google and Bing, and it has aucun Google or Bing equivalent. Its objectif, per Yandex: “The Yandex robot uses ce directive to éviter reloading duplicate information. Ce improves the robot’s efficiently and reduces le serveur charger.” (Que “efficiently” is a genuine typo on Yandex’s live page — I’m quoting it as-is plutôt que silently fixing it.)
The problem it solves: “The nouveau parameter que doesn’t affecter lune page content may result in duplicate pages que ne doit pas be inclus in the search.” The syntax:
Clean-param: p0[&p1&p2&..&pn] [path]Yandex’s propre worked exemple — three URLs que differ seulement by a ref tracking
parameter:
www.example.com/some_dir/get_book.pl?ref=site_1&book_id=123
www.example.com/some_dir/get_book.pl?ref=site_2&book_id=123
www.example.com/some_dir/get_book.pl?ref=site_3&book_id=123…collapse to un URL canonique (www.example.com/some_dir/get_book.pl?book_id=123)
with a unique directive:
User-agent: Yandex
Clean-param: ref /some_dir/get_book.plTwo details faire Clean-param facile to obtenir incorrect. Premier, “The Clean-param directive
ne fait pas exiger mandatory combination with the Disallow directive” — it stands on
its propre; vous don’t Disallow the parameter URLs. Second, it’s intersectional:
per Yandex it “is intersectional, so it peut be specified anywhere in the fichier,
regardless of the emplacement.” Unlike Allow/Disallow, qui are anchored to a
chemin, Clean-param is a global directive vous pouvez drop anywhere in the fichier.
Yandex aussi notes it may handle some parameters automatically: “Parameters pour analytics and tracking que don’t affecter lune page content may be automatically supprimé by the moteur de recherche si the algorithms determine que ceux parameters are insignificant.” But relying on Clean-param pour the ones que matter to vous is the deterministic déplacer. Ce is the Yandex analog to how Google now leans on canonicalization signals (balise canonical, maillage interne) since it deprecated its old URL Parameters outil — Yandex simplement donne vous an explicit robots.txt directive où Google doesn’t.
Crawl-delay is dead (since Feb 2018)
Si you’ve lire que Crawl-delay fonctionne in Yandex’s robots.txt, que information is
stale. Yandex’s propre dedicated page is unambiguous: “From February 22, 2018, Yandex
doesn’t prendre into account the Crawl-delay directive.”
Ce is worth flagging parce que au moins un widely-read SEO resource encore dit the
opposite. Ahrefs’ robots.txt guide (authored by Joshua Hardwick, pas me) currently
states “Google no longer supports this directive, but Bing and Yandex do.” On
the Yandex half, that’s contradicted by Yandex’s propre current documentation. I’d
treat Yandex’s dedicated, dated page as authoritative ici — but the broader lesson
is the utile un: vérifier a crawl-behavior claim contre the engine’s live docs
avant vous trust a secondhand guide, parce que ces details drift and même bon
sources go stale. (Remarque the contrast with Bing, qui fait encore honor
crawl-delay — un of the réel Bingbot/YandexBot divergences.)
The replacement is the Fréquence d’exploration setting in Yandex Webmaster, qui lets vous influence how fast YandexBot récupère votre site. (Yandex’s Crawl-delay page is essentially un sentence plus a pointer to que setting.)
How YandexBot handles JavaScript
Yandex’s JavaScript rendering is explicitly labeled beta (β) by Yandex itself, and the par défaut behavior is “at the bot’s discretion” — the bot “va independently determine si to execute JavaScript code on le site’s pages.” Quand it fait, it may “assess the quality and completeness of le contenu on the pages with and sans JavaScript” and serve whichever version is probable plus utile to the visitor.
There’s a meaningful tension worth surfacing. The 2023 Yandex source-code leak (exposed internal engineering docs, pas an official statement — I’ll cover it plus ci-dessous) suggested a simpler picture. As Mike King wrote in his Moteur de recherche Land analysis of the leak: “Yandex has aucun separate rendering system pour JavaScript. Ils dire ce in leur documentation and, although ils have Webdriver-based system pour visual regression testing appelé Gemini, ils limite themselves to text-based explorer.” (Que internal “Gemini” is a Yandex visual-regression testing outil — complètement unrelated to Google’s Gemini AI model, despite the shared nom. Worth disambiguating so nobody conflates the two.) And Dan Taylor’s Moteur de recherche Journal write-up concluded there’s “nothing nouveau to suggest Yandex peut explorer JavaScript yet outside of déjà publicly documented processes.”
So the honest réponse is neither “YandexBot renders JS just like Googlebot” nor “YandexBot never touches JavaScript.” It’s selective and beta, by Yandex’s propre description. The leak’s “no separate rendering system” account is consistent with que framing, but it’s leaked internal material relayed via industry coverage, pas a Yandex disclosure — so treat “architecturally simpler than Google” as a plausible reading of two consistent signals, pas a documented fact. Don’t prendre soit account on faith pour a route que matters to vous; tester it directement:
- Rendered output: récupérer lune page with JavaScript disabled and comparer it to the JS-rendered version. Si the two differ meaningfully, don’t assume Yandex saw the rendered un.
- Resource accès: confirmer the JS, CSS, and API endpoints lune page dépend on
aren’t blocked in
robots.txt. Yandex’sYandexRenderResourcesBotrécupère render-time resources, but (per Yandex’s propre documentation of it) seulement pour pages the principal indexation bot peut déjà reach — a blocked resource on an allowed page encore won’t charger. - Delayed and interactive content: anything que loads après the
DOMContentLoadedevent or behind a click isn’t guaranteed to render. Yandex’s avancé rendering settings (window.YandexRotorSettings) exist specifically pour sites où “content loads with a delay” — that’s a signal worth reading, pas simplement a config option.
The practical takeaway: Yandex’s propre docs recommend vous “Prohibit rendering si SSR (Rendu côté serveur) or pre-rendering is implemented on le site,” and remarque que “Executing JavaScript code may create additional load on your server.” Si vous vouloir reliable Yandex indexation of JS-heavy content, serve it server-side plutôt que testing votre luck on rendu côté client.
Pour AJAX-style sites, Yandex dit “Quand indexation an AJAX site, the Yandex bot scans
the original URLs and executes JavaScript code on les” — and it has déplacé away
from the old HTML-snapshot hack: si vous encore utiliser the deprecated
meta name="fragment" approach, “the bot will ignore it and index the original page.”
Its modern recommendation mirrors Google’s: “Si the liens on AJAX pages utiliser the #
character, modifier the addresses to URLs sans ce character. Par exemple, vous may
utiliser the History API.”
Ce que the 2023 source-code leak revealed à propos de the robot d’exploration
In January 2023, Yandex’s internal code source leaked — a well-corroborated event covered à travers Moteur de recherche Land, Moteur de recherche Journal, and others. Treat ce as leaked internal documentation, pas a Yandex statement, but it exposed réel detail à propos de how the robot d’exploration fonctionne. Per Mike King’s SEL analysis: “Yandex’s documentation discusses a dual-distributed robot d’exploration system. Un pour real-time exploration appelé the ‘Orange Robot d’exploration’ and un autre pour general exploration.” He drew a parallel to Google, qui “is said to have had an index stratified into three buckets, un pour housing real-time explorer, un pour regularly crawled and un pour rarely crawled.” Les deux engines, En d’autres termes, apparaître to embrace segmented exploration driven by how souvent content updates.
The leak aussi tied exploration directement to architecture du site. Per Dan Taylor’s SEJ coverage, “URLs que are reachable from the homepage have a ‘plus élevé’ level of importance.” That’s a clean bridge from “robot d’exploration mechanics” to “pourquoi internal linking matters” — the même explorer depth logic que s’applique to Googlebot.
(Pour scale: coverage noted the widely-cited “1,922 ranking factors” figure was spécifique to un archive fichier, with the fuller codebase reportedly containing far plus à travers multiple fichiers — garder a date and a source on quelconque spécifique number si vous cite un.)
Pourquoi YandexBot encore matters — and how courant it is in robots.txt
Yandex’s global share is tiny, but its Russia share n’est pas: ~71% in Russia vs. Google’s ~27% as of June 2026 (StatCounter). Que stability is the entier raison YandexBot deserves separate treatment. Si vous have quelconque Russia/CIS-facing business, blocking YandexBot forecloses the dominant engine in que market.
How souvent do sites même bother configuring pour it? Rarely, but rising. In the Web Almanac 2022 SEO chapter (I was a reviewer que année; I was lead author of the 2021 chapter): YandexBot appeared in “simplement 0,5% of robots.txt fichiers in 2021. By 2022, là was a six-fold augmenter, with 3% of fichiers specifying Yandexbot.” Petit, but a clair upward trend — and a utile “how common is this in the wild” baseline.
Si you’re deciding si to block it, the honest framing is a business question, pas a technical un — and it’s a réel debate que plays out in webmaster forums. Pour the broader mechanics YandexBot lives à l’intérieur — URL discovery, the explorer scheduler, rendering, and the crawl-vs-index-vs-rank distinctions — voir the exploration hub. And pour configuring Yandex specifically as partie of a Russia/CIS strategy, the international-SEO and market-specific-SEO material ties it ensemble.
AI summary
A condensed prendre on the Avancé version:
- YandexBot = Yandex Search’s principal indexation robot d’exploration — the Yandex equivalent of
Googlebot/Bingbot. UA string:
Mozilla/5,0 (compatible; YandexBot/3,0; +http://yandex.com/bots) AppleWebKit/537,36 (KHTML, comme Gecko) Chrome/81.0.4044.268(don’t match on the Chrome version — it changements). - Two robots.txt tokens, différent scopes:
YandexBot= principal indexation bot seulement;Yandex= the broader bot family. AUser-agent: Yandexblock rend Yandex ignoreUser-agent: *. Some Yandex bots may ignore robots.txt entirely. - Clean-param is Yandex-only — consolidates URL parameters que don’t modifier
content (aucun Google/Bing equivalent). It doesn’t exiger
Disallow, and it’s intersectional (peut go anywhere in the fichier). - Crawl-delay is dead — Yandex stopped honoring it on February 22, 2018; utiliser the Fréquence d’exploration outil à la place. Some SEO guides encore wrongly claim Yandex supports it. (Bing fait; Yandex doesn’t.)
Disallow≠noindex— disallowed pages “can participate in Yandex search.” Utilisernoindex(meta or HTTP header) to en réalité supprimer une page.- JavaScript rendering is beta and “at the bot’s discretion.” The 2023 leak suggested “no separate rendering system for JavaScript” — so it’s selective and simpler que Google’s two-wave pipeline. Préférer SSR/pre-rendering; éviter hash-bang URLs (utiliser the History API).
- 2023 source-code leak (leaked, pas official): revealed a dual-crawler system (real-time “Orange Crawler” + general robot d’exploration) and que homepage-reachable URLs carry plus élevé importance (crawl-depth as a signal).
- Vérifier with reverse-then-forward DNS to a
yandex.ru/yandex.net/yandex.comhost — the même FCrDNS technique Google and Bing utiliser — jamais the spoofable user-agent. - Pourquoi care: ~71% Russia search share (vs. Google ~27%, June 2026). Tiny globally, dominant regionally. In robots.txt: ~0,5% of fichiers (2021) → 3% (2022) per the Web Almanac.
Documentation officielle
Primary-source documentation, mostly from Yandex Webmaster Aider, with Google/Bing contrast liens. Heads up: several Yandex Webmaster outil pages (Fréquence d’exploration, the robots.txt analyzer) are app pages que besoin a logged-in navigateur to voir entièrement, and Google’s/Bing’s verification pages are JavaScript-rendered — confirmer details live si a lien bounces.
Yandex
- The User-agent directive — the
YandexBotvs.Yandextoken scoping and precedence rules. - Comment vérifier que a robot belongs to Yandex — the complet user-agent string and the reverse-DNS verification méthode.
- En utilisant robots.txt — the directives Yandex recognizes, the 500 KB / HTTP 200 fichier requirements, and the
Disallow≠ index remarque. - The Clean-param directive — Yandex’s unique parameter-consolidation directive (syntax + worked exemple).
- The Crawl-delay directive — lune page stating Crawl-delay has been ignored since February 22, 2018.
- Site fréquence d’exploration — the Crawl-delay replacement in Yandex Webmaster.
- Indexation pages with JavaScript (β) — the “at the bot’s discretion” par défaut and the SSR recommendation.
- Indexation AJAX sites — JS execution on original URLs and the History-API recommendation.
Google / Bing (pour contrast)
- Vérifier Requêtes from Google Robots d’exploration and Fetchers — Google’s reverse/forward-DNS méthode (contre
googlebot.com/google.com/googleusercontent.com) — the même FCrDNS pattern Yandex uses. - Qui robots d’exploration fait Bing utiliser? — Bing’s robot d’exploration liste; Bing verifies contre
*.search.msn.comand encore honorscrawl-delay(Yandex doesn’t).
Quotes from the source
On-the-record statements from Yandex’s propre documentation, plus the leaked-internal material relayed via industry coverage. Chaque lien is a deep lien que jumps to the quoted passage. Yandex’s docs are English translations publié by Yandex itself — quoted verbatim from ceux pages.
Yandex — robots.txt tokens and precedence
- “If the
User-agent: Yandexstring is detected, theUser-agent: *string is ignored.” Jump to quote - “Some Yandex robots may ignore directives in
robots.txt, including those forUser-agent: Yandex.” Jump to quote - “
YandexAdditionalBot… Helps process robots.txt to prevent page content from appearing in Search with Yandex AI responses. Applies to the pages that have been indexed by the primary crawler.” Jump to quote
Yandex — verifying the bot
- “Some robots can disguise themselves as Yandex robots by indicating the relevant User Agent. You can check the authenticity of a robot using a reverse DNS lookup.” Jump to quote
- “Check whether the host belongs to Yandex. All Yandex robots have names ending in
yandex.ru,yandex.netoryandex.com.” Jump to quote - “If the IP addresses do not match, it means that the host name is fake.” Jump to quote
Yandex — robots.txt and Clean-param
- “Pages restricted in
robots.txtcan participate in Yandex search. To remove pages from search, specify thenoindexdirective in the HTML code of the page or configure the HTTP header.” Jump to quote - “The Yandex robot uses this directive to avoid reloading duplicate information. This improves the robot’s efficiently and reduces the server load.” [sic — “efficiently” is Yandex’s propre typo] Jump to quote
- “The Clean-param directive does not require mandatory combination with the Disallow directive.” Jump to quote
- “Do not restrict such pages in
robots.txt, or the Yandex bot can’t index them and detect your instructions.” Jump to quote
Yandex — Crawl-delay and JavaScript
- “From February 22, 2018, Yandex doesn’t take into account the Crawl-delay directive.” Jump to quote
- “Prohibit rendering if SSR (Server-Side Rendering) or pre-rendering is implemented on the site.” Jump to quote
- “When indexing an AJAX site, the Yandex bot scans the original URLs and executes JavaScript code on them.” Jump to quote
The 2023 Yandex source-code leak (leaked internal documentation, relayed via industry coverage — pas an official Yandex statement)
- “Yandex has no separate rendering system for JavaScript. They say this in their documentation and, although they have Webdriver-based system for visual regression testing called Gemini, they limit themselves to text-based crawl.” — Mike King, iPullRank, via Moteur de recherche Land. Lire the coverage
- “Yandex’s documentation discusses a dual-distributed crawler system. One for real-time crawling called the ‘Orange Crawler’ and another for general crawling.” — Moteur de recherche Land. Lire the coverage
- “There’s nothing new to suggest Yandex can crawl JavaScript yet outside of already publicly documented processes.” — Dan Taylor, via Moteur de recherche Journal. Lire the coverage
Web Almanac 2022, SEO chapter (I was a reviewer)
- “Yandexbot was specified in just 0.5% of robots.txt files in 2021. By 2022, there was a six-fold increase, with 3% of files specifying Yandexbot.” Jump to quote
Devrait I block, autoriser, or throttle YandexBot?
The YandexBot question is almost toujours a business question wearing a technical costume. Fonctionner via it.
What to do about YandexBot in your logs
YandexBot Mythes et erreurs à éviter
The traps que come up la plupart souvent — several are widely-repeated and worth correcting:
- “Crawl-delay slows YandexBot down.” Aucun. Yandex stopped honoring
Crawl-delayon February 22, 2018 (“Yandex doesn’t prendre into account the Crawl-delay directive”). Some SEO guides encore claim Yandex supports it — that’s stale. Utiliser the Fréquence d’exploration setting in Yandex Webmaster à la place. (Do-instead: vérifier crawl-behavior claims contre Yandex’s live docs, and throttle via Fréquence d’exploration, pas Crawl-delay.) - “
Disallowin robots.txt removes the page from Yandex.” Aucun — Yandex’s propre docs dire disallowed pages “can participate in Yandex search.”Disallowarrête exploration, pas indexation. (Do-instead: utilisernoindexin the HTML or an HTTP header to en réalité exclude une page; let Yandex explorer it so it peut voir the tag.) - “
User-agent: YandexBotandUser-agent: Yandexare the same thing.” Aucun —YandexBottargets seulement the principal indexation bot;Yandextargets the broader family. And aUser-agent: Yandexblock rend Yandex ignore votreUser-agent: *block entirely. (Do-instead: pick the token que matches the scope vous en réalité intend, and don’t assume*rules reach Yandex bots.) - “YandexBot renders JavaScript just like Googlebot.” Pas established. Yandex’s rendering is beta and “at the bot’s discretion,” and the 2023 leak suggested “aucun separate rendering system pour JavaScript.” (Do-instead: serve JS-dependent content via SSR/pre-rendering plutôt que assuming rendu côté client va be indexé.)
- “The user-agent string proves it’s YandexBot.” Aucun — it’s trivially spoofed.
(Do-instead: vérifier with reverse-then-forward DNS to a
yandex.ru/yandex.net/yandex.comhost, exactly as you’d vérifier Googlebot or Bingbot.) - “Any YandexBot traffic on a non-Russian site is inherently suspicious/fake.”
Pas necessarily — legitimate YandexBot crawls globally-facing sites aussi, pas simplement
.rudomains. (Do-instead: vérifier by DNS avant deciding it’s spoofed; a réel Yandex host is réel regardless of votre audience.) - “The 2023 leak proved Yandex has a secret superior JS crawler.” Aucun — coverage concluded there’s “nothing nouveau to suggest Yandex peut explorer JavaScript yet outside of déjà publicly documented processes.” (Do-instead: don’t over-read the leak; it corroborated the limited-rendering picture, it didn’t overturn it.)
YandexBot — cheat sheet
User-agent string (principal indexation bot)
Mozilla/5.0 (compatible; YandexBot/3.0; +http://yandex.com/bots) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/81.0.4044.268Match on the YandexBot token, pas the Chrome/81.0.4044.268 version — Yandex
dit le navigateur version may modifier.
The two robots.txt tokens
| Token | Scope |
|---|---|
User-agent: YandexBot | The principal indexation bot seulement |
User-agent: Yandex | Yandex’s broader bot family |
A User-agent: Yandex block causes Yandex to ignore votre User-agent: * block.
Some Yandex bots may ignore robots.txt entirely.
Directives Yandex recognizes
| Directive | Ce que it fait |
|---|---|
User-agent | Qui robot the rules appliquer to |
Disallow | Prohibits exploration (pas indexation) |
Allow | Permet exploration/indexation of sections or pages |
Sitemap | Chemin to le sitemap fichier |
Clean-param | Yandex-only — ignore listed URL parameters (aucun Google/Bing equivalent) |
Clean-param rapide formulaire
User-agent: Yandex
Clean-param: ref /some_dir/get_book.plConsolidates ?ref=… variants to un URL. Doesn’t besoin Disallow; intersectional
(peut go anywhere in the fichier).
Vérifier vs. Google vs. Bing (tout utiliser forward-confirmed reverse DNS)
| Engine | Reverse-DNS host doit fin in |
|---|---|
| Yandex | yandex.ru / yandex.net / yandex.com |
googlebot.com / google.com / googleusercontent.com | |
| Bing | *.search.msn.com |
Fast facts
- Crawl-delay: dead since Feb 22, 2018 — utiliser the Fréquence d’exploration outil. (Bing
encore honors
crawl-delay; Yandex doesn’t.) robots.txtdoit be namedrobots.txt, be ≤ 500 KB, and retourner HTTP 200.- JavaScript rendering: beta, “at the bot’s discretion” — préférer SSR.
- Russia share: ~71% (vs. Google ~27%), June 2026 (StatCounter).
- In robots.txt: 0,5% (2021) → 3% (2022) of fichiers (Web Almanac).
Vérifier a bot is really YandexBot
The user-agent is trivially spoofed, so confirmer with a reverse-DNS lookup (doit fin
in a Yandex domain) followed by a forward-DNS lookup (doit resolve back to the même
IP). Ce is the exact FCrDNS pattern you’d utiliser pour Googlebot or Bingbot — seulement the
domain suffixes differ (yandex.ru / yandex.net / yandex.com).
macOS / Linux
# 1) Reverse-DNS the IP from your logs — the host must end in
# yandex.ru, yandex.net, or yandex.com
host 5.255.253.1
# → 1.253.255.5.in-addr.arpa domain name pointer <something>.yandex.com (illustrative)
# 2) Forward-DNS that hostname back — it must resolve to the same IP
host <the-hostname-from-step-1>.yandex.comWindows
nslookup 5.255.253.1
nslookup <the-hostname-from-step-1>.yandex.comSi the reverse lookup doesn’t fin in a Yandex domain, or the forward lookup doesn’t retourner the original IP, it isn’t YandexBot — drop or rate-limit it.
Extract verified-vs-spoofed YandexBot hits from a log fichier
A rapide shell réussir to pull “YandexBot” lines and vérifier chaque source IP’s PTR. Adjust field positions pour votre log format (ce assumes a courant combined format with the IP premier).
# Pull unique IPs that claimed to be YandexBot, then reverse-resolve each
grep -i 'YandexBot' access.log \
| awk '{print $1}' | sort -u \
| while read ip; do
host="$(host "$ip" 2>/dev/null | awk '/pointer/{print $NF}' | sed 's/\.$//')"
case "$host" in
*.yandex.ru|*.yandex.net|*.yandex.com) echo "REAL $ip $host" ;;
"" ) echo "NO-PTR $ip" ;;
* ) echo "FAKE $ip $host" ;;
esac
doneREAL lines encore deserve a forward-lookup confirmation pour anything you’ll act on
(host "$host" devrait retourner the original IP), but ce triages the obvious
impostors premier.
Regex to match the YandexBot token in a user-agent
Match the product token, pas the Chrome version (qui changements). Case-insensitive:
YandexBot/\d+(\.\d+)?Broader “any Yandex robot” match (catches YandexImages, YandexMobileBot, etc.):
Yandex[A-Za-z]*/\dRemember: a matching user-agent is necessary but pas sufficient — pair quelconque regex match with the DNS vérifier ci-dessus avant trusting it.
A robots.txt starting point pour Yandex
# Slow/duplicate-parameter cleanup for all Yandex bots
User-agent: Yandex
Clean-param: utm_source&utm_medium&utm_campaign /
# Block only the main indexing bot from a low-value space
User-agent: YandexBot
Disallow: /internal-search/
Sitemap: https://example.com/sitemap.xmlRemember: Disallow blocks exploration, pas indexation — utiliser noindex to supprimer a
page from Yandex search. To entièrement block tout Yandex bots (some ignore robots.txt),
block by verified IP at le serveur.
YandexBot readiness checklist
A rapide réussir to confirmer you’re handling YandexBot deliberately, pas by accident:
- You’ve decided si Yandex matters pour votre site (quelconque Russia/CIS audience or business?) — the whole block/autoriser question hinges on ce.
- Bot verification uses reverse + forward DNS to a
yandex.ru/yandex.net/yandex.comhost — pas the spoofable user-agent, and pas a hardcoded IP liste. - You’re en utilisant the correct robots.txt token pour votre intent:
YandexBot(principal indexation bot seulement) vs.Yandex(the broader family). - Vous know a
User-agent: Yandexblock rend Yandex ignore votreUser-agent: *block. - You’re pas relying on
Crawl-delay(dead since Feb 22, 2018) — throttle via the Fréquence d’exploration outil in Yandex Webmaster à la place. - You’re pas en utilisant
Disallowto deindex — that’snoindex’s job (with exploration allowed). - Duplicate URL parameters are consolidated with Clean-param où relevant
(it doesn’t besoin
Disallow, and it peut go anywhere in the fichier). - JS-dependent content is served via SSR/pre-rendering, since Yandex’s rendering is beta and “at the bot’s discretion.”
- AJAX routes utiliser clean URLs / the History API, pas hash-bang (
#!) URLs. - Pour complet exclusion of tout Yandex bots (some ignore robots.txt), you’re blocking by verified IP at le serveur, pas simplement in robots.txt.
Outils pour verifying and monitoring YandexBot
The two jobs que come up la plupart with YandexBot are confirming a hit is réel and seeing how beaucoup it’s en réalité exploration vous. Commencer ici:
- Googlebot Verifier — paste an IP from votre logs and it runs
the forward-confirmed reverse-DNS vérifier pour vous (the même méthode Yandex documents
pour confirming a
yandex.ru/yandex.net/yandex.comhost), naming the réel network owner quand it’s a spoofer à la place. Faster que runninghost/nslookupby hand pour every suspicious hit, and it aussi covers Googlebot, Bingbot, and the major AI robots d’exploration si you’re checking a mixed log. - Log Fichier Analyzer — drop in a server accès log (nginx, Apache, IIS/W3C, or JSON) and voir YandexBot’s réel explorer footprint: how beaucoup of votre site it’s hitting, qui sections, status-code waste, and a spoofer report flagging IPs que claim to be YandexBot but don’t vérifier out. Utile pour the “is this crawling actually straining my server” question from the decision tree ci-dessus — everything runs in votre navigateur, nothing is uploaded.
From Yandex itself
- Yandex Webmaster — the Fréquence d’exploration setting (the Crawl-delay replacement) and the robots.txt analyzer live ici; you’ll besoin a verified, logged-in Yandex Webmaster account to utiliser soit.
Ressources utiles
My connexe writing
- Indexé, though blocked by robots.txt — pourquoi a robots-blocked URL encore obtient indexé; the même
Disallow≠noindextrap s’applique to Yandex. - Robots.txt and SEO: Everything Vous devez Know — the general robots.txt référence (remarque: its Yandex
Crawl-delayclaim is outdated, per Yandex’s propre dated docs — a bon exemple of verifying contre the source). - The Story of Blocking 2 High-Ranking Pages With Robots.txt — my first-party experiment on ce que en réalité se produit quand vous block ranking pages.
- Meet the Nouveau Web Robots d’exploration: AI Bots Are Closing in on Moteur de recherche Bots — how the robot d’exploration cast in votre logs has modifié.
- The SEO Bots Que ~140 Million Websites Block the La plupart — my study with Xibeijia Guan on robots.txt block rates (it covers Western SEO-tool bots, pas YandexBot — cited pour the methodology and as a contrast to how rarely sites configurer pour Yandex).
My speaking
- How Search Fonctionne (SlideShare) — my walkthrough of exploration, rendering, indexation, and ranking. (Standing disclaimer s’applique: “This is my understanding of systems… not going to be 100% complete or accurate.”)
From autour the industry
- Yandex scrapes Google and autre SEO learnings from le code source leak (Mike King / iPullRank, Moteur de recherche Land, Jan 30, 2023) — the “Orange Crawler” dual-crawler system and “no separate rendering system for JavaScript” findings.
- Yandex Données Leak: The Ranking Factors & The Myths We Trouvé (Dan Taylor, Moteur de recherche Journal, Feb 1, 2023) — the crawl-depth-as-importance signal and the “nothing new on JavaScript crawling” conclusion.
- Yandex ‘leak’ reveals 1 922 search ranking factors (Moteur de recherche Land) — context and dating pour the leak’s scale.
- Fait the Yandex Code Leak Tell Us Anything À propos de Google? (seoClarity) — a mesuré cross-engine lire of the leak.
- The Ultimate Guide to Yandex SEO (Moteur de recherche Journal) — broader Yandex-optimization context au-delà the robot d’exploration.
- Web Almanac 2022 — SEO chapter (HTTP Archive) — the source of the YandexBot-in-robots.txt adoption stat.
Stats worth citing
- ~71% Russia search share — Yandex’s share of the Russian search market versus Google’s ~27%, per StatCounter (données reported pour June 2026; time-sensitive, so cite the month/année and expect drift). Ce is the entier raison YandexBot warrants separate treatment from “block all non-Google bots.” Source
- 0,5% → 3% of robots.txt fichiers — YandexBot went from being specified in “simplement 0,5% of robots.txt fichiers in 2021” to “3% of fichiers specifying Yandexbot” in 2022 — a six-fold augmenter, though encore petit (Web Almanac 2022 SEO chapter, qui I reviewed). Utile as a historical baseline pour “how common is this in the wild.” Source
- February 22, 2018 — the date Yandex stopped honoring
Crawl-delay, per its propre documentation. A precise, citable deprecation date to counter stale guides claiming Yandex encore supports the directive. Source
Testez vos connaissances: YandexBot
Five rapide questions on YandexBot, its robots.txt behavior, and how it compares to Googlebot and Bingbot. Pick an réponse pour chaque, alors vérifier.
Journal des modifications
Mis à jour le 18 juil. 2026.
Résumé éditorial et détails enregistrés des changements.Détails des changements
-
Les notes détaillées des changements sont actuellement disponibles en anglais.
-
Les notes détaillées des changements sont actuellement disponibles en anglais.
-
Les notes détaillées des changements sont actuellement disponibles en anglais.
-
Les notes détaillées des changements sont actuellement disponibles en anglais.
-
Les notes détaillées des changements sont actuellement disponibles en anglais.
Comparaison complète indisponible — aucun instantané antérieur n’a été archivé pour cette révision.