Baidu SEO
How to rank in Baidu, mainland China's dominant search engine — and why Google rules don't transfer. Hosting behind the Great Firewall, ICP filing (beian), Simplified Chinese, Baidu's own-property SERPs, the meta tags it supports, the HTTPS and MIP timelines, and the 2026 market-share reality.
Baidu SEO means making a site accessible and understandable to Baidu Search. Validate crawlable delivery in the target market, localize content for users, follow applicable hosting rules, and use Baidu's official Search Resource Platform for verification, diagnostics, and URL submission. Treat market-share estimates, account gates, and historical crawler claims as volatile context rather than ranking rules.
Evidence for this claim Baidu's official Search Resource Platform provides webmaster guidance and site-management tools for Baidu Search. Scope: Current official Baidu webmaster platform; no market-share claim. Confidence: high · Verified: Baidu Search Resource Platform Evidence for this claim Baidu provides official link-submission workflows, but submission is not an indexing or ranking guarantee. Scope: Current Baidu link-submission interface. Confidence: medium · Verified: Baidu Search Resource Platform: Link submissionTL;DR — Baidu SEOBaidu SEO is the practice of optimizing a website to rank in Baidu, mainland China's dominant search engine. It differs structurally from Google SEO: hosting location behind the Great Firewall, an ICP filing to legally host on mainland servers, Simplified Chinese content, Baidu's preference for its own properties, and its own crawler (Baiduspider) and webmaster console (Ziyuan) all matter more than any translated Google playbook. is the work of making a site accessible and understandable to Baidu Search. It is not just “Google SEO translated into Chinese”: confirm reliable access for your target users and Baidu’s crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index., localize the content, follow applicable hosting rules, and use Baidu’s own Search Resource Platform for diagnostics and URL submission.
What Baidu SEO is
Google isn’t the search engine everywhere. In mainland China, the leader is Baidu — and Google itself hasn’t run a normal, uncensored search product there since 2010. So if your audience is in China, the Google playbook you already know mostly doesn’t apply. That’s what Baidu SEO is: optimizing a site to rank in Baidu, on Baidu’s terms.
The reason it’s so different from Google SEO is that the hardest parts aren’t about your content at all. They’re about plumbing and paperwork:
- Where your site lives. A site hosted far from China loads slowly and unreliably behind the Great Firewall — the government system that filters and slows traffic in and out of the country. Slow, flaky pages are bad for crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor., indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed., and users.
- The ICP filing (beian). To legally host a site on a server or CDN inside mainland China, you need a government registration called an ICP filing. It’s a legal/hosting gate, not something Baidu invented — but it’s often the first real hurdle.
- The language. Baidu’s home market is mainland China, which uses Simplified Chinese. Traditional Chinese (Hong Kong, Taiwan, Macau) is a different market where Baidu barely competes.
Baidu likes its own websites best
Here’s a difference you’ll notice immediately in the results: Baidu tends to put its own properties at the top — its encyclopedia (Baidu Baike), its Q&A site (Baidu Zhidao), its blogging arm (Baijiahao), and its forums (Tieba). So part of “ranking on Baidu” is often getting your brand into those Baidu properties, not just onto your own pages.
The things people get wrong
- “Baidu still has ~70–80% of the market.” Not anymore. It’s real, but it’s been shrinking — Bing and other Chinese engines have taken a big chunk.
- “I just need to translate my site.” Translation is the easy part. Hosting, reachability, and the ICP filing are the parts that actually block sites.
- “You must host in China to rank.” You don’t have to, but hosting outside China usually means slower, less reliable access — which hurts.
Want the full technical version — the exact ICP rules, the meta tagsMeta tags are HTML elements in a page's head that pass metadata about the page to search engines and browsers. For SEO only a few matter — the title element, the meta description, and the robots meta tag — while meta keywords and most others are ignored. Baidu supports, the HTTPSHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.' and mobile (MIP) history, why JavaScript is a bigger problem on Baidu, and the current market-share numbers? Switch to the Advanced tab.
Baidu SEO launch checklist
Market and infrastructure
- Confirm mainland China is an addressable market before creating a separate program.
- Test real page and asset performance from behind the Great Firewall.
- If using mainland hosting or a mainland CDN, complete the required ICP filing process; do not describe it as a direct ranking factor.
- Remove critical dependencies on third-party resources that are blocked or unreliable in mainland China.
Content and rendering
- Publish useful Simplified Chinese content reviewed for local terminology and intent.
- Render primary content and links in server-delivered HTML rather than requiring client-side JavaScript.
- Use crawlable internal linksAn internal link is a hyperlink from one page on a website to another page on the same website. Internal links help search engines discover your pages and pass ranking signals (PageRank and anchor-text context) between them. and a stable URL structureURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly..
- Treat Baidu-owned properties and the ad-heavy result layout as part of the competitive SERP, not as a reason to copy Google tactics.
Technical discovery
- Verify BaiduspiderBaiduspider is the web crawler operated by Baidu, China's dominant search engine — the Chinese-market equivalent of Googlebot. It discovers, crawls, and indexes pages for Baidu Search, and you manage its access via robots.txt and Baidu's own console, Baidu Search Resource Platform (Ziyuan). requests by reverse and forward DNS before changing crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. controls.
- Review robots rules, response status, canonical signals, and sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. coverage for the Chinese URLs.
- Configure relevant Baidu-supported metadata intentionally; do not expect hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. to provide Baidu targeting.
- Register and monitor the site in Baidu Search Resource PlatformBaidu Webmaster Tools — officially the Baidu Search Resource Platform (百度搜索资源平台) at ziyuan.baidu.com — is Baidu's free webmaster console: verify a site, submit URLs and sitemaps (including the fast API push), and monitor indexing and crawl diagnostics. Since May 2022 registration needs a Chinese mobile number..
Ongoing QA
- Recheck China-side availability after template, CDN, analytics, or consent changes.
- Inspect logs for verified BaiduspiderBaiduspider is the web crawler operated by Baidu, China's dominant search engine — the Chinese-market equivalent of Googlebot. It discovers, crawls, and indexes pages for Baidu Search, and you manage its access via robots.txt and Baidu's own console, Baidu Search Resource Platform (Ziyuan). crawl patterns and recurring error responses.
- Review rankings and result features using current market data rather than repeating an old market-share assumption.
- Revalidate content quality and compliance with local owners before expanding page volume.
Evidence for this claim Baidu's official Search Resource Platform provides webmaster guidance and site-management tools for Baidu Search. Scope: Current official Baidu webmaster platform; no market-share claim. Confidence: high · Verified: Baidu Search Resource Platform Evidence for this claim Baidu provides official link-submission workflows, but submission is not an indexing or ranking guarantee. Scope: Current Baidu link-submission interface. Confidence: medium · Verified: Baidu Search Resource Platform: Link submissionTL;DR — Baidu SEOBaidu SEO is the practice of optimizing a website to rank in Baidu, mainland China's dominant search engine. It differs structurally from Google SEO: hosting location behind the Great Firewall, an ICP filing to legally host on mainland servers, Simplified Chinese content, Baidu's preference for its own properties, and its own crawler (Baiduspider) and webmaster console (Ziyuan) all matter more than any translated Google playbook. requires engine-specific validation rather than importing every Google assumption. Prioritize crawlable server-rendered output, localized content, reliable delivery in the target market, and applicable hosting rules. Use Baidu’s official platform to verify current crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index., submission, and diagnostic behavior; market share and account requirements are volatile and are not ranking factors.
Baidu SEO is a different engine, not a translated Google
This site’s market-specific SEO hub frames the core of it already, and this is the deep dive it points to: getting indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. well in China “is as much about where your site is hosted and whether it’s written in simplified Chinese as it is about content.” I’ll operationalize that here. The mental trap to avoid is treating Baidu as Google-with-a-language-swap. What broadly does transfer from your Google work: content quality, site speed as a factor, backlinks as a signal, and structured markup where supported. What does not transfer: hosting/ICP realities, Great Firewall reachability, hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. (Baidu doesn’t support it), a SERP dominated by the engine’s own assets, and a mobile framework (MIP) that never existed on Google’s side.
Baidu itself defines the discipline plainly in its official Baidu SEO Guide 2.0: 搜索引擎优化…指为了提升网页在搜索引擎自然搜索结果中(非商业性推广结果)的收录数量以及排序位置而做的优化行为 — roughly, SEO is optimization to improve a page’s inclusion count and ranking position in a search engine’s natural results, excluding commercial promotion results. That “excluding commercial promotion results” clause matters on Baidu specifically, because its SERPs are famously ad-heavy — the organic game and the paid game are visibly separate surfaces.
Hosting location and speed behind the Great Firewall
The single biggest infrastructural difference from Google SEO is that where your server sits is a first-order concern. A site hosted in the US or EU is reachable from mainland China, but traffic crossing the Great Firewall is slow and inconsistent — which degrades crawlabilityCrawlability is how well search engine crawlers can discover, access, and fetch a site's pages. A crawlability issue is any technical condition — blocked access, broken links, server failures, or bloated URL inventory — that stops pages from reaching the index., indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. speed, and real user experience, all of which feed rankings indirectly. Hosting in or near mainland China (or on a China-optimized CDN) is the practical fix.
There’s a concrete gotcha worth flagging from this site’s Hexo SEOHexo SEO is the practice of optimizing sites built with Hexo, the Node.js static site generator. Hexo outputs pure static HTML at build time, so content is crawlable on the first fetch — but sitemaps, robots.txt, meta descriptions, canonicals, and structured data all depend on plugins and theme configuration. guide: GitHub Pages is blocked in mainland China, so a site published there is effectively invisible to Chinese users — use Gitee Pages, Cloudflare, or a Chinese CDN for Baidu traffic. The same logic applies to any embedded third-party resource that lives on a blocked host: Google Fonts, Google Analytics, YouTube embeds, and many Western CDNs either fail or hang behind the Firewall, so audit your page’s third-party requests and replace blocked ones with China-reachable equivalents.
ICP filing (beian) — what it actually is, and what it isn’t
This is the most-misunderstood piece of China SEO, so be precise about it. An ICP filing (备案, bèi’àn) is a PRC government (MIIT) requirement, not a Baidu-issued ranking rule. It was established by the Telecommunications Regulations of the People’s Republic of China (中华人民共和国电信条例), promulgated in September 2000, per Wikipedia’s ICP-license entry: “This license regime was instated by the Telecommunications Regulations of the People’s Republic of China (中华人民共和国电信条例) that was promulgated in September 2000.” The enforcement teeth: “China-based Internet service providers are required to block the site if a license is not acquired within a grace period.”
It comes in two tiers, and Cloudflare’s China Network ICP docs describe them in the cleanest practitioner language I’ve foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore.:
- ICP filing (备案 / “Bei’An”) — the first level: “An ICP filing enables the holder to host a website on a server or CDN in Mainland China for informational purposes only” (non-commercial / non-transactional).
- ICP license (ICP许可证 / “ICP Zheng”) — the commercial tier: it “allows online platforms or third-party sellers selling goods and services to deploy their website on a hosting server or CDN within Mainland China.”
- And the sequencing: “Companies acquiring an ICP license must already have obtained an ICP filing” first.
Here’s the nuance almost every competitor guide gets wrong. ICP is a hosting/legal gate, not a documented Baidu ranking factor. You need it to legally host on a mainland server or CDN and to get fast, reliable access — but Baidu has never confirmed it as a direct algorithmic signal. The data backs the distinction: Search Engine Journal’s 2023 Baidu ranking-factors data study by Marcus Pentzek found “less than half (48%) of the top-ranking pages have an ICP reference,” which the study says is “corroborated by our experience with client websites without licenses still achieving rankings.” So frame ICP honestly: get one to legally host in mainland China and load fast there, but don’t repeat the common myth that Baidu won’t rank you without it.
Language and market scope: Simplified vs. Traditional Chinese
Baidu’s market is mainland China, which uses Simplified Chinese. That’s the language you write and optimize in for Baidu. Traditional Chinese serves Hong Kong, Taiwan, and Macau — markets where Baidu has minimal share and Google dominates. Conflating “Chinese-language SEO” with “Baidu SEO” is a scoping error: if you only have Traditional Chinese content aimed at Taiwan, you’re mostly doing Google SEO, not Baidu SEO. And because Baidu keys off hosting location and content-language rather than markup, the hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. signal you’d lean on for Google does nothing here — as this site’s hreflang guideHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. puts it, “Baidu doesn’t support hreflang at all (it keys off hosting location, Chinese domain registration, ICP licensing, and content-language).” See international keyword research for confirming which engine your audience actually uses before you spend on any of this.
Baidu’s preference for its own properties (and the ad-heavy SERP)
Expect the top organic real estate to be occupied by Baidu itself. Per Search Engine Land’s May 2026 “fragmented search ecosystem” piece (Marcus Pentzek): “The most coveted organic positions are almost always reserved for Baidu’s own properties,” naming “Baidu Baike (the encyclopedia), Baidu Zhidao (the Q&A hub), and Baijiahao (the news/blogging arm)” as “the permanent residents of Page 1.” Dragon Metrics’ Baidu SEO guide independently corroborates with a number: “Baidu’s own properties can occupy up to 70% of real estate on page 1,” and “for almost all search queries, you can see at least one Baidu property ranking in the top 5 positions.” And the SERP is ad-dense — per the same SEL piece, “It isn’t uncommon to see ads claiming the top, middle, and bottom of a Baidu search engine results page (SERP).”
The practical implication: a chunk of Baidu visibility work is placement inside Baidu’s ecosystem — a well-maintained Baidu Baike entry, answers on Zhidao, a Baijiahao account, a Tieba presence — not just ranking your own domain.
Technical setup for Baidu
Baidu Search Resource Platform (Ziyuan)
Baidu’s webmaster console is the Baidu Search Resource Platform (百度搜索资源平台, Ziyuan) — the direct analogue to Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. and Bing Webmaster ToolsMicrosoft's free portal for monitoring and improving how a site appears in Bing search — the peer to Google Search Console, plus IndexNow instant indexing, richer backlink data, and keyword volumes. Because Bing's index also feeds Microsoft Copilot, it doubles as a window into AI-search visibility.: site verification, sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing. and link submission, an index-volume tool, and site-data dashboards. It’s almost entirely in Chinese, with no English equivalent to Google’s Search Central. The access hurdles that trip up non-Chinese teams (a Chinese mobile number, real-name gates, the API push mechanics) are their own topic — see the dedicated Baidu Webmaster Tools deep dive rather than solving them here.
Meta tags Baidu supports
A persistent myth is that Baidu doesn’t honor meta robots directives. It does, and
it also documents its own Baiduspider-named variants — verbatim from Baidu’s own
“Methods to prevent search engine indexing”
page:
- Standard nofollowrel=\"nofollow\" is a value of the HTML link rel attribute that tells search engines you don't vouch for a linked page and don't want to pass ranking signals to it. Since 2019–2020 Google treats it as a hint, not a directive — and it does not reliably block crawling or indexing.: “如果您不想搜索引擎追踪此网页上的链接,且不传递链接的权重,请将此元标记置入网页的<head>部分:
<meta name="robots" content="nofollow">” — don’t follow links on this page or pass their weight. - Baidu-only nofollow: “要允许其他搜索引擎跟踪,但仅防止百度跟踪您网页的链接…
<meta name="Baiduspider" content="nofollow">” — let other engines follow, block only Baidu. - Standard noarchiveNoarchive is a robots meta directive that tells search engines not to keep a cached copy of a page. Google's current documentation lists it under historical and unused rules and says it's ignored, because the cached-link feature it once controlled no longer exists — but on Bing it still controls whether content is used in Bing Chat/Copilot answers and AI training. It is not an access-control measure.: “要防止所有搜索引擎显示您网站的快照…
<meta name="robots" content="noarchive">” — no cached snapshot in any engine. - Baidu-only noarchive: “要允许其他搜索引擎显示快照,但仅防止百度显示…
<meta name="Baiduspider" content="noarchive">” — block only Baidu’s snapshot.
Frame these correctly: the Baiduspider-named tag is a parallel mechanic to
Google’s own engine-specific <meta name="googlebot">, not an exotic Chinese
invention. The real difference is that Baidu documents and expects engine-specific
tags more routinely. As for meta keywords: don’t re-litigate it here — this
site’s meta keywords guideThe meta keywords tag — <meta name=\"keywords\" content=\"...\"> — is a mid-1990s HTML head element meant to let a page declare its own topic keywords to search engines. Google has publicly ignored it for ranking since 2009, and no major search engine uses it as a ranking signal today. It's dead as SEO, and a populated one only leaks your target keywords to anyone who views your source. already
covers that “Baidu’s guidance has flip-flopped: dismissed as obsolete in a 2012
statement, then revived with relevance language in a 2020 update.” (More on that
2012 claim, and why it’s shakier than it looks, in the Anti-patterns tab.)
HTTPS on Baidu
Baidu supports HTTPSHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.', and has for years. Per Search Engine Journal (Dan Taylor, Oct 2019): “Baidu announced full support for crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. and indexing HTTPS protocol webpages in 2015, and then in summer 2016 they further update Baidu-spider to better handle HTTPS.” That’s roughly a year after Google’s own August-2014 HTTPS-as-ranking-signal announcement — a clean timeline comparison. There was also a tooling milestone: per Hermes Ma’s 2017 Search Engine Land piece (verified via archive), “Baidu Webmaster Tools launched a new feature of HTTPS Site Authentication in May” (2017). And adoption has climbed — per SEJ’s 2023 data study, “the adoption of HTTPS has risen from 55% in 2020 (Searchmetrics’ study) to 69.6% among top-ranking pages.”
Mobile: MIP’s rise (2016) and Cache shutdown (2020)
This is the date to get right, because it’s where most competitor content is stale. Baidu’s AMPAMP (Accelerated Mobile Pages) is an open-source web framework Google launched in 2015 to make mobile pages load near-instantly via restricted HTML/CSS/JS and CDN caching. It was never a ranking factor and, since June 2021, is no longer required for Top Stories.-equivalent was MIP (Mobile Instant Pages), launched mid-August 2016 — Dragon Metrics reported it about a week before its August 22, 2016 post: “About a week ago Baidu announced their own version of AMP – the Mobile Instant Pages (MIP),” adding that “The technical aspects of MIP is almost the same as AMP, which isn’t a surprise as it’s been a usual practice for Baidu to follow the footsteps of Google.”
MIP is now effectively dead. Baidu’s own MIP Cache service discontinuation notice, dated April 24, 2020, states the MIP Cache service would be discontinued — 公示 through May 31, 2020, with the MIP entry in the platform closed and the Cache service phased out over June 1–30, 2020. One nuance worth keeping: the notice clarifies “MIP核心、组件等前端静态资源仍然会正常维护与使用” — MIP’s core components and front-end static resources would keep being maintained; it was the caching/CDN layer that was retired, not MIP markup being deindexed overnight. But the caching product is what made MIP meaningfully different from a normal fast page, and it’s been gone since mid-2020. Six years on, any guide still describing MIP as current best practice is stale — build fast, standards-based mobile pages instead.
JavaScript rendering limits
BaiduspiderBaiduspider is the web crawler operated by Baidu, China's dominant search engine — the Chinese-market equivalent of Googlebot. It discovers, crawls, and indexes pages for Baidu Search, and you manage its access via robots.txt and Baidu's own console, Baidu Search Resource Platform (Ziyuan). handles JavaScript poorly — a genuine, checkable difference from modern GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer., which does render JS (with delay and queueing). This site’s Hexo SEOHexo SEO is the practice of optimizing sites built with Hexo, the Node.js static site generator. Hexo outputs pure static HTML at build time, so content is crawlable on the first fetch — but sitemaps, robots.txt, meta descriptions, canonicals, and structured data all depend on plugins and theme configuration. guide states it directly: “Baidu’s crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. handles JavaScript poorly, so Hexo’s static HTML is a direct advantage here.” The takeaway generalizes: for Baidu specifically, serve important content, links, and metadata in the initial server-rendered or static HTML — don’t depend on client-side renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. to inject what needs to be indexed. If you’re on a specific stack, the BaiduspiderBaiduspider is the web crawler operated by Baidu, China's dominant search engine — the Chinese-market equivalent of Googlebot. It discovers, crawls, and indexes pages for Baidu Search, and you manage its access via robots.txt and Baidu's own console, Baidu Search Resource Platform (Ziyuan). deep dive and the Hexo guide cover the crawler-level and plugin-level details.
My Render Gap Checker compares initial HTML with a Chromium-rendered result. It is not a Baiduspider emulator or a mainland-China reachability test, so use it only to identify client dependence and confirm the outcome in Baidu’s own tools.
Test representative templates with my Render Gap Checker and move essential content, links, and metadata into the initial response when the rendered version carries the page. Render Gap Checker Free
- Compare the initial and rendered content on each critical template.
- Server-render any essential text, links, titles, canonicals, and supported metadata.
- Validate Baiduspider access and indexing in Baidu Search Resource Platform from the target market.
Baidu vs. Google vs. Bing in China: the 2026 market-share picture
Kill the “~70–80% share” cliché — it’s the single most common staleness error in Baidu content. Baidu’s dominance is real but shrinking, and the market is genuinely fragmenting. Two dated StatCounter snapshots from within a single quarter show both the level and the volatility (all devices, China):
| Month (StatCounter, China) | Baidu | Bing | Haosou/360 | Yandex | Sogou | |
|---|---|---|---|---|---|---|
| April 2026 | ~44.62% | ~22.55% | ~18.27% | ~10.35% | ~2.08% | ~1.88% |
| June 2026 | 47.44% | 23.24% | 14.02% | 10.23% | 2.62% | 2.28% |
Source: StatCounter, China. Two things jump out. First, Baidu sits in the mid-40s, not the 70–80% of legacy guides. Second, the swing between two months — Baidu +2.8 points, Haosou −4.25 points — is itself the proof that this number is volatile; re-pull it at any serious decision point rather than trusting a static figure. And note Bing: ~23% is a strong #2, well above the “negligible in China” claim you’ll still see in older content. As the SEL May-2026 piece frames it, Baidu SEO isn’t dead — your website just isn’t the sole focus of Chinese web search anymore, which is why a Baidu-plus-Bing-plus-Haosou view increasingly beats a Baidu-only one.
Where this fits
This is one of the market-specific SEO deep dives, alongside the Yandex (Russia) and Naver (South Korea) guides and the Japan piece. The crawler and console mechanics have their own homes: BaiduspiderBaiduspider is the web crawler operated by Baidu, China's dominant search engine — the Chinese-market equivalent of Googlebot. It discovers, crawls, and indexes pages for Baidu Search, and you manage its access via robots.txt and Baidu's own console, Baidu Search Resource Platform (Ziyuan). for the bot, and the Baidu Webmaster Tools deep dive for Ziyuan. For the wider strategy of going international — choosing markets, languages, and URL structuresURL structure is how the parts of a web address — scheme, domain, path, query string, and fragment — are organized and formatted. It mostly affects crawling, usability, and how engines understand a page, not rankings directly. — see the International SEOInternational SEO is the practice of optimizing a site so search engines understand which countries and/or languages it targets, and serve the right version to each user. It spans URL structure, hreflang, and on-page localization. pillar.
AI summary
A condensed take on the Advanced version:
- Baidu SEOBaidu SEO is the practice of optimizing a website to rank in Baidu, mainland China's dominant search engine. It differs structurally from Google SEO: hosting location behind the Great Firewall, an ICP filing to legally host on mainland servers, Simplified Chinese content, Baidu's preference for its own properties, and its own crawler (Baiduspider) and webmaster console (Ziyuan) all matter more than any translated Google playbook. ≠ Google SEO in Chinese. The hardest problems are infrastructural and legal, not on-page. What transfers: content quality, speed, backlinks, supported structured markup. What doesn’t: hosting/ICP, Great Firewall reachability, hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others., own-property SERPs, MIP.
- Hosting matters most. Host in/near mainland China (or a China CDN) so pages load fast behind the Great Firewall. GitHub Pages is blocked; audit third-party resources (Google Fonts/Analytics, YouTube) that fail behind the Firewall.
- ICP filing (beian) is a PRC government requirement to host on mainland servers/CDNs — not a confirmed Baidu ranking factor. SEJ foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. only 48% of top-ranking pages carry a visible ICP reference; sites without one still rank. Two tiers: filing (informational) → license (commercial), filing first.
- Simplified Chinese for mainland/Baidu; Traditional (HK/Taiwan/Macau) is a Google market. Baidu ignores hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others., keying off hosting + content-language.
- Baidu favors its own properties (Baike, Zhidao, Baijiahao, Tieba) in top slots, and SERPs are ad-heavy — so ecosystem placement is part of the work.
- Meta tagsMeta tags are HTML elements in a page's head that pass metadata about the page to search engines and browsers. For SEO only a few matter — the title element, the meta description, and the robots meta tag — while meta keywords and most others are ignored.: Baidu supports standard
robotsplusBaiduspider-namednofollow/noarchive(a parallel to Google’sgooglebottag). - HTTPSHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.' supported since 2015–2016 (69.6% adoption among top pages by 2023). MIP (AMPAMP (Accelerated Mobile Pages) is an open-source web framework Google launched in 2015 to make mobile pages load near-instantly via restricted HTML/CSS/JS and CDN caching. It was never a ranking factor and, since June 2021, is no longer required for Top Stories. clone, launched 2016) is effectively dead — Cache service retired April 24, 2020.
- BaiduspiderBaiduspider is the web crawler operated by Baidu, China's dominant search engine — the Chinese-market equivalent of Googlebot. It discovers, crawls, and indexes pages for Baidu Search, and you manage its access via robots.txt and Baidu's own console, Baidu Search Resource Platform (Ziyuan). renders JS poorly — serve server-rendered/static HTML.
- Market share is volatile: ~44.62% (Apr 2026) → 47.44% (Jun 2026), far below the historical 60–80%; Bing is a strong ~23% #2. Re-pull figures at decision time.
Official documentation
Baidu’s webmaster documentation is almost entirely in Chinese; there is no English-language equivalent to Google’s Search Central. Primary sources:
Baidu (百度搜索资源平台 / Ziyuan)
- Baidu Search Resource Platform (home) — verification, sitemapA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing./link submission, indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.-volume tool, site data.
- Baidu SEO Guide 2.0 (百度搜索引擎优化指南2.0) — Baidu’s own SEO handbook; the foreword carries its SEO definition.
- Methods to prevent search engine indexing (禁止搜索引擎收录的方法) — the
robotsandBaiduspider-specific meta-tag directives. - Introduction to Baiduspider (百度spider介绍) — crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. definition, per-vertical user-agentsA user agent is the HTTP request header a client (browser, crawler, or bot) sends to identify itself. For crawlers, a short user-agent token — a substring of that string — is what robots.txt rules actually target., and reverse-DNS verification.
- MIP Cache service discontinuation notice (MIP Cache 服务下线通知) — dated April 24, 2020; the MIP-cachingCaching stores a copy of a page or resource — in a browser, a CDN edge node, or a search crawler's own cache — so it can be served again without regenerating or re-downloading it. It isn't a direct ranking factor, but it feeds page speed and crawl efficiency. shutdown.
- Mobile Search Website Building Optimization White Paper (百度移动搜索建站优化白皮书) — last updated June 20, 2018.
Government / legal (ICP)
- ICP license — Wikipedia — legal basis (Telecommunications Regulations of the PRC, September 2000).
- Cloudflare China Network — ICP — practitioner-facing breakdown of the filing vs. license tiers.
Google/Bing intersection
- Google Search Central — International SEO — the hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. model that doesn’t apply to Baidu.
- Google — “A new approach to China” (Jan 12, 2010) — the post announcing Google’s exit from uncensored mainland search.
Quotes from the source
Baidu has no English-language rep culture comparable to Google’s Search Liaison or Bing’s Fabrice Canel — so the “on the record” material here is Baidu’s own written docs (in Chinese, with English paraphrase — a translation is never presented as a verbatim English quote) plus the primary Google and legal sources.
Baidu — its own definition of SEO (Baidu SEOBaidu SEO is the practice of optimizing a website to rank in Baidu, mainland China's dominant search engine. It differs structurally from Google SEO: hosting location behind the Great Firewall, an ICP filing to legally host on mainland servers, Simplified Chinese content, Baidu's preference for its own properties, and its own crawler (Baiduspider) and webmaster console (Ziyuan) all matter more than any translated Google playbook. Guide 2.0, foreword)
- “搜索引擎优化…指为了提升网页在搜索引擎自然搜索结果中(非商业性推广结果)的收录数量以及排序位置而做的优化行为” — SEO is optimization to improve a page’s inclusion count and ranking position in a search engine’s natural results, excluding commercial promotion results. Jump to the source
Baidu — the meta tagsMeta tags are HTML elements in a page's head that pass metadata about the page to search engines and browsers. For SEO only a few matter — the title element, the meta description, and the robots meta tag — while meta keywords and most others are ignored. it supports (Methods to prevent search engine indexingStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed.)
- “要允许其他搜索引擎跟踪,但仅防止百度跟踪您网页的链接…
<meta name="Baiduspider" content="nofollow">” — the Baidu-onlynofollowvariant: let other engines follow links, block only Baidu. Jump to the source - “要允许其他搜索引擎显示快照,但仅防止百度显示…
<meta name="Baiduspider" content="noarchive">” — the Baidu-onlynoarchive: block only Baidu’s cached snapshot. Jump to the source
Baidu — MIP Cache shutdown (dated April 24, 2020)
- “MIP核心、组件等前端静态资源仍然会正常维护与使用” — MIP’s core components and front-end static resources will keep being maintained; i.e., the cachingCaching stores a copy of a page or resource — in a browser, a CDN edge node, or a search crawler's own cache — so it can be served again without regenerating or re-downloading it. It isn't a direct ranking factor, but it feeds page speed and crawl efficiency. layer was retired, not MIP markup being deindexed. Jump to the source
Google — leaving uncensored mainland search (A new approach to China, Jan 12, 2010)
- “we are no longer willing to continue censoring our results on Google.cn, and so over the next few weeks we will be discussing with the Chinese government the basis on which we could operate an unfiltered search engine within the law, if at all.” Jump to quote
Legal — the ICP regime (Wikipedia, “ICP license”)
- “China-based Internet service providers are required to block the site if a license is not acquired within a grace period.” Jump to quote
Cloudflare — the two ICP tiers (China Network docs)
- “An ICP filing enables the holder to host a website on a server or CDN in Mainland China for informational purposes only.” Jump to quote
Note: the ziyuan.baidu.com pages are Chinese-language and partly JavaScript-rendered, and the Google 2010 post and some source pages resist automated fetching (verified via archive) — confirm each quote against the live page, and treat the English alongside each Chinese quote as paraphrase, not a verbatim English quotation.
Do you actually need Baidu SEO — and what first?
Most teams asking “can we just also rank on Baidu?” haven’t decided whether the China audience justifies the infrastructural lift. Walk the branch before spending.
Scoping a Baidu SEO effort
Baidu vs. Google at a glance
| Dimension | Baidu | |
|---|---|---|
| Hosting location | Largely irrelevant to ranking | First-order — host in/near China, fast behind the Great Firewall |
| Legal gate | None | ICP filing (beian) to host on mainland servers (PRC law, not a ranking rule) |
| Language | Any, with hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. | Simplified Chinese (mainland); Traditional = a Google market |
| hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. | Supported | Not supported — keys off hosting + content-language |
| JavaScript renderingTurning HTML, CSS, and JavaScript into the final visual page and DOM. | Renders JS (delayed) | Poor — favor server-rendered/static HTML |
| Own-property SERP bias | Modest | Heavy — Baike, Zhidao, Baijiahao, Tieba fill top slots |
| Ad density in SERP | Moderate | High — ads top, middle, and bottom |
| Fast-mobile framework | AMPAMP (Accelerated Mobile Pages) is an open-source web framework Google launched in 2015 to make mobile pages load near-instantly via restricted HTML/CSS/JS and CDN caching. It was never a ranking factor and, since June 2021, is no longer required for Top Stories. (deprecating) | MIP — effectively dead (Cache retired Apr 24, 2020) |
| HTTPSHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.' | Ranking signal since 2014 | Supported since 2015–2016 |
| Webmaster console | Google Search ConsoleA free Google service that reports how a site performs in Google Search and surfaces problems with how Google crawls, indexes, and serves it. It's first-party data straight from Google — but you don't need it to appear in results. | Baidu Search Resource PlatformBaidu Webmaster Tools — officially the Baidu Search Resource Platform (百度搜索资源平台) at ziyuan.baidu.com — is Baidu's free webmaster console: verify a site, submit URLs and sitemaps (including the fast API push), and monitor indexing and crawl diagnostics. Since May 2022 registration needs a Chinese mobile number. (Ziyuan), Chinese-only |
| CrawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. | GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. | BaiduspiderBaiduspider is the web crawler operated by Baidu, China's dominant search engine — the Chinese-market equivalent of Googlebot. It discovers, crawls, and indexes pages for Baidu Search, and you manage its access via robots.txt and Baidu's own console, Baidu Search Resource Platform (Ziyuan). |
Meta tagsMeta tags are HTML elements in a page's head that pass metadata about the page to search engines and browsers. For SEO only a few matter — the title element, the meta description, and the robots meta tag — while meta keywords and most others are ignored. Baidu documents
| Tag | Effect |
|---|---|
<meta name="robots" content="nofollow"> | Don’t follow links / pass weight (all engines) |
<meta name="Baiduspider" content="nofollow"> | Block only Baidu from following links |
<meta name="robots" content="noarchive"> | No cached snapshot in any engine |
<meta name="Baiduspider" content="noarchive"> | Block only Baidu’s snapshot |
Fast facts
- Market share (StatCounter, China): ~44.62% (Apr 2026) → 47.44% (Jun 2026) — not the historical 60–80%. Re-pull at decision time.
- Bing in China: ~23% — a strong #2, not “negligible.”
- ICP reference among top-ranking pages: 48% (SEJ, 2023) — ICP is a gate, not a confirmed ranking factor.
- HTTPSHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.' adoption among top pages: 69.6% (SEJ, 2023, up from 55% in 2020).
- MIP lifecycle: launched mid-Aug 2016, Cache service retired Apr 24, 2020.
Baidu SEO myths that cost you
Myth: “You must host in mainland China to rank at all on Baidu.” Why it’s wrong: hosting outside China won’t get you banned or algorithmically penalized by name — SEJ’s 2023 study states “any website accessible in China can rank.” The real effect is indirect: offshore hosting is slower and less reliable through the Great Firewall, which hurts crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. and UX. Do instead: host in/near China or on a China-optimized CDN for speed, and get an ICP filing only because you need it to legally host on mainland servers — not because Baidu won’t rank you without one.
Myth: “ICP is a Baidu ranking requirement.” Why it’s wrong: ICP is a PRC government (MIIT) legal/hosting requirement, and SEJ’s data study foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. only 48% of top-ranking pages carry a visible ICP reference — with client sites lacking licenses “still achieving rankings.” Do instead: treat ICP as a legal gate for mainland hosting and reachability, and stop repeating “no ICP, no ranking.”
Myth: “MIP is how you build fast Baidu mobile pages.” Why it’s wrong: Baidu retired the MIP Cache service on April 24, 2020 — six years stale. Only MIP’s front-end components were kept “maintained and used”; the cachingCaching stores a copy of a page or resource — in a browser, a CDN edge node, or a search crawler's own cache — so it can be served again without regenerating or re-downloading it. It isn't a direct ranking factor, but it feeds page speed and crawl efficiency. layer that made MIP distinct is gone. Do instead: build fast, standards-based mobile pages; don’t invest in MIP as a differentiator.
Myth: “Baidu doesn’t support meta robotsThe robots meta tag is an HTML element in a page's head — <meta name=\"robots\" content=\"noindex\"> — that tells search engines how to index and serve that page. It's crawl-then-obey: a page blocked in robots.txt is never fetched, so the tag is never seen. tags.”
Why it’s wrong: Baidu’s own docs document standard robots (nofollow,
noarchive) and Baiduspider-specific variants.
Do instead: use the standard tags, and reach for the Baiduspider-named variants
only when you want to control Baidu specifically without affecting other engines.
Myth: “Baidu still commands ~70–80% of Chinese search.” Why it’s wrong: StatCounter put Baidu at ~44.62% (Apr 2026) and 47.44% (Jun 2026) amid a fragmenting market, with Bing around 23%. Do instead: cite share with a date, plan for a multi-engine reality (Baidu + Bing + Haosou), and re-pull the numbers before any real decision.
Myth: “A Baidu engineer named ‘Lee’ officially killed meta keywordsThe meta keywords tag — <meta name=\"keywords\" content=\"...\"> — is a mid-1990s HTML head element meant to let a page declare its own topic keywords to search engines. Google has publicly ignored it for ranking since 2009, and no major search engine uses it as a ranking signal today. It's dead as SEO, and a populated one only leaks your target keywords to anyone who views your source. in August 2012 — it’s a documented fact.” Why it’s wrong: this widely-repeated claim traces to a single unlinked industry post (Jinray’s China SEO Diary), and even Search Engine Journal cites “spokesperson Lee” secondhand with no date, quote, or link. I couldn’t locate any primary Baidu source for it. It’s an unverifiable-but-repeated SEO claim, not confirmed fact. Do instead: rely on this site’s hedged meta keywordsThe meta keywords tag — <meta name=\"keywords\" content=\"...\"> — is a mid-1990s HTML head element meant to let a page declare its own topic keywords to search engines. Google has publicly ignored it for ranking since 2009, and no major search engine uses it as a ranking signal today. It's dead as SEO, and a populated one only leaks your target keywords to anyone who views your source. framing (a “2012 statement,” a 2020 revival) rather than asserting a named engineer and exact date.
Myth: “Google SEO knowledge transfers directly to Baidu.” Why it’s wrong: it conflates two ecosystems. Hosting/ICP, Great Firewall reachability, hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. (unsupported), own-property SERPs, and MIP do not transfer. Do instead: carry over what does — content quality, speed, backlinks, supported structured markup — and treat everything else as a different engine.
Verify a bot really is Baiduspider
The Baiduspider user-agentA user agent is the HTTP request header a client (browser, crawler, or bot) sends to identify itself. For crawlers, a short user-agent token — a substring of that string — is what robots.txt rules actually target. is trivially spoofed. Baidu’s own FAQ says a genuine
BaiduspiderBaiduspider is the web crawler operated by Baidu, China's dominant search engine — the Chinese-market equivalent of Googlebot. It discovers, crawls, and indexes pages for Baidu Search, and you manage its access via robots.txt and Baidu's own console, Baidu Search Resource Platform (Ziyuan). hostname is formatted as *.baidu.com or *.baidu.jp — anything else
is an impersonator. Confirm with a reverse + forward DNS check.
macOS / Linux
# 1) Reverse-DNS the IP from your logs — a real Baiduspider hostname ends in
# .baidu.com or .baidu.jp
host 180.76.15.30
# → ... domain name pointer baiduspider-180-76-15-30.crawl.baidu.com
# 2) Forward-DNS that hostname back — it must resolve to the same IP
host baiduspider-180-76-15-30.crawl.baidu.comWindows
nslookup 180.76.15.30
nslookup baiduspider-180-76-15-30.crawl.baidu.comIf the reverse lookup doesn’t end in .baidu.com / .baidu.jp, or the forward
lookup doesn’t match the original IP, it isn’t BaiduspiderBaiduspider is the web crawler operated by Baidu, China's dominant search engine — the Chinese-market equivalent of Googlebot. It discovers, crawls, and indexes pages for Baidu Search, and you manage its access via robots.txt and Baidu's own console, Baidu Search Resource Platform (Ziyuan)..
Find Baiduspider hits in your access log (regex)
Separate the real crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. by user-agent, then confirm by DNS above. The
per-vertical variants share the Baiduspider tokenA token is the smallest unit of text (or image/audio/video) an LLM processes — roughly 4 characters, or about ¾ of an English word. A context window is the maximum number of tokens (input plus output) a model can hold at once, like its short-term memory. family:
# All Baiduspider variants (web, image, video, news, mobile) in an access log
grep -E 'Baiduspider(-image|-video|-news)?' /var/log/nginx/access.log
# Just the requests, counted by path — spot what Baidu is actually crawling
grep -Ei 'baiduspider' access.log \
| awk '{print $7}' | sort | uniq -c | sort -rn | head -20Audit a page for Firewall-blocked third-party resources (DevTools Console)
Paste into Chrome DevTools’ Console on the page you’re auditing to list every external host it loads — then check which are blocked in mainland China (Google Fonts/Analytics, YouTube, many Western CDNs, GitHub Pages):
// Unique third-party hosts this page requests
[
...new Set(
performance
.getEntriesByType("resource")
.map((r) => new URL(r.name).host)
.filter((h) => h !== location.host),
),
]
.sort()
.forEach((h) => console.log(h));Bookmarklet: quick Simplified-vs-Traditional and lang check
Drag this to your bookmarks bar; it flags the declared lang and whether the page
looks like it’s targeting Baidu’s mainland (Simplified) market:
javascript: (() => {
const l = document.documentElement.lang || "(none)";
const ok = /^zh(-|_)?(cn|hans|CN|Hans)?$/i.test(l);
alert(
"html lang: " +
l +
"\nLooks like Simplified/mainland target: " +
(ok ? "yes" : "no — Baidu wants zh-CN / zh-Hans"),
);
})();Example IPs and hostnames above are illustrative — always verify against the actual IPs in your own logs; don’t hardcode any single Baidu IP range.
Resources worth your time
My related writing and this site’s Baidu-relevant guides
- Market-Specific SEO — the hub this deep dive sits under (Baidu, Yandex, Naver, Japan).
- hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. — why Baidu ignores hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others. and keys off hosting + content-language instead.
- Meta keywordsThe meta keywords tag — <meta name=\"keywords\" content=\"...\"> — is a mid-1990s HTML head element meant to let a page declare its own topic keywords to search engines. Google has publicly ignored it for ranking since 2009, and no major search engine uses it as a ranking signal today. It's dead as SEO, and a populated one only leaks your target keywords to anyone who views your source. — the flip-flop history (2012 statement → 2020 revival), so this article doesn’t re-derive it.
- Hexo SEOHexo SEO is the practice of optimizing sites built with Hexo, the Node.js static site generator. Hexo outputs pure static HTML at build time, so content is crawlable on the first fetch — but sitemaps, robots.txt, meta descriptions, canonicals, and structured data all depend on plugins and theme configuration. — Baidu-specific setup on a static-site stack (GitHub Pages is blocked in China; static HTML helps Baidu’s weak JS crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor.).
- International SEO keyword research — confirm which engine your audience actually uses first.
- International SEO: The Weird Technical Parts (Pubcon Vegas 2019) — my deck on international SEOInternational SEO is the practice of optimizing a site so search engines understand which countries and/or languages it targets, and serve the right version to each user. It spans URL structure, hreflang, and on-page localization.’s technical traps (hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others., sitemapsA sitemap is a file that lists the pages, images, videos, and other files on your site so search engines can discover them. It helps discovery, but submitting a sitemap doesn't guarantee crawling or indexing.). It’s Google/Yandex-focused, not Baidu — a general credibility anchor, not a China source.
From around the industry
- How China’s fragmented search ecosystem is reshaping SEO in 2026 (Search Engine Land, Marcus Pentzek) — the own-property SERP dominance and market-fragmentation framing.
- Baidu Ranking Factors for 2024: A Comprehensive Data Study (Search Engine Journal, Marcus Pentzek) — the 48%-ICP-reference and HTTPSHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.'-adoption data.
- Baidu SEO: Content Delivery, Speed & Accessibility (Search Engine Journal, Dan Taylor) — the HTTPSHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.' 2015/2016 timeline.
- What SEOs need to know about Baidu in 2017 (Search Engine Land, Hermes Ma) — the May-2017 HTTPS Site Authentication launch and MIP-over-AMPAMP (Accelerated Mobile Pages) is an open-source web framework Google launched in 2015 to make mobile pages load near-instantly via restricted HTML/CSS/JS and CDN caching. It was never a ranking factor and, since June 2021, is no longer required for Top Stories. recommendation.
- The Ultimate Guide to Baidu SEO in China (Dragon Metrics) — the “up to 70% of page-1 real estate” own-property stat.
- Baidu Launches Their Own Version of AMP – MIP (Dragon Metrics) — the mid-August-2016 MIP launch.
- Do you need an ICP license to rank on Baidu? (Chinafy) — practitioner take on the ICP-as-ranking-factor question.
- StatCounter — Search Engine Market Share, China — the live, dated share figures (re-pull before citing).
Test yourself: Baidu SEO
Five quick questions on optimizing for China’s dominant search engine. Pick an answer for each, then check.
Baidu SEO
Baidu SEO is the practice of optimizing a website to rank in Baidu, mainland China's dominant search engine. It differs structurally from Google SEO: hosting location behind the Great Firewall, an ICP filing to legally host on mainland servers, Simplified Chinese content, Baidu's preference for its own properties, and its own crawler (Baiduspider) and webmaster console (Ziyuan) all matter more than any translated Google playbook.
Related: Baiduspider, Baidu Webmaster Tools (Ziyuan)
Baidu SEO
Baidu SEO is optimizing a site to rank in Baidu (百度), the search engine that leads inside mainland China — a market where Google has not run an uncensored product since 2010. Baidu’s own guide defines SEO as optimization to improve a page’s inclusion count and ranking position in a search engine’s natural results, excluding commercial promotion results (非商业性推广结果).
What makes it structurally different from Google SEO is that the hardest problems are infrastructural and legal, not on-page. A site generally needs to be hosted in or near mainland China and load fast behind the Great Firewall; to legally host on a mainland server or CDN it needs an ICP filing (beian, 备案) — a PRC government requirement, not a Baidu ranking rule. Content should be in Simplified Chinese (mainland), not Traditional (which serves Hong Kong, Taiwan, and Macau, where Baidu has little share). Baidu runs its own crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index., BaiduspiderBaiduspider is the web crawler operated by Baidu, China's dominant search engine — the Chinese-market equivalent of Googlebot. It discovers, crawls, and indexes pages for Baidu Search, and you manage its access via robots.txt and Baidu's own console, Baidu Search Resource Platform (Ziyuan)., and its own console, Baidu Search Resource PlatformBaidu Webmaster Tools — officially the Baidu Search Resource Platform (百度搜索资源平台) at ziyuan.baidu.com — is Baidu's free webmaster console: verify a site, submit URLs and sitemaps (including the fast API push), and monitor indexing and crawl diagnostics. Since May 2022 registration needs a Chinese mobile number. (Ziyuan). It does not support hreflangHreflang is an annotation (in HTML, HTTP headers, or XML sitemaps) that tells search engines which language and optional region a page targets, and which alternate versions exist. It only works when every page in the cluster references all the others., keys off hosting and content-language instead, handles JavaScript poorly, and fills its top SERP positions with its own properties — Baidu Baike, Zhidao, Baijiahao, and Tieba.
Getting the details right matters: HTTPSHTTPS is the encrypted version of HTTP — it uses TLS to authenticate the server and protect data in transit between a browser and a website. Google announced it as a lightweight ranking signal in 2014 and today conditionally prefers HTTPS pages as canonical; Chrome marks plain HTTP pages 'Not Secure.' has been supported since 2015–2016, Baidu’s AMPAMP (Accelerated Mobile Pages) is an open-source web framework Google launched in 2015 to make mobile pages load near-instantly via restricted HTML/CSS/JS and CDN caching. It was never a ranking factor and, since June 2021, is no longer required for Top Stories.-equivalent MIP is effectively dead (the MIP Cache service was retired in 2020), and Baidu’s ~45–47% share in 2026 is a far cry from its historical 60–80% dominance as Bing and Haosou gain ground. Baidu SEO is a different engine, not a translated Google strategy.
Related: Baiduspider, Baidu Webmaster Tools (Ziyuan)
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.