Agentic Search
Agentic search is AI that autonomously plans, browses, and synthesizes a multi-step research task — not one answer to one query. How it differs from AI search.
1 evidence signal on this page
- Related live toolSEO MCP Server
Agentic search is AI that autonomously browses, researches, and synthesizes over minutes — not just answers a single query. The key technical gotcha: agentic browsers spin up real Chrome/Edge VMs and can't be blocked via robots.txt.
TL;DR — Agentic searchAgentic search is an AI system that autonomously plans, runs, and synthesizes a multi-step research or browsing task — dozens to hundreds of searches and real page reads — instead of returning one answer to a single query. All agentic search is AI search; not all AI search is agentic. is AI that does the whole research job for someone, not just answers one question. Instead of a quick reply, it goes off and runs dozens of searches, reads the actual pages, and comes back with a written-up answer — sometimes minutes later. Think ChatGPT giving you a fast answer vs. ChatGPT Deep Research disappearing for ten minutes and returning a report. For you, the practical thing to know: some of these agents browse as a real Chrome browser, so
robots.txtcan’t stop them.
The three layers of search
“Agentic search” is an evolving product pattern rather than a single standardized ranking architecture. Evidence for this claim Deep-research products can perform multi-step web research and synthesize cited results. Scope: Documented product capabilities, not a universal definition or ranking architecture for agentic search. Confidence: high · Verified: OpenAI: Introducing deep research Documented systems may browse, synthesize, and take multi-step actions, but behavior varies by product and release. Evidence for this claim The term agentic search covers an evolving set of vendor implementations whose tools and behavior can differ. Scope: Editorial synthesis from documented product behavior; implementations and releases may vary. Confidence: medium · Verified: OpenAI: Introducing deep research
It helps to think of search as having grown three layers, one on top of the next:
- Traditional search. You type a query, Google hands back a list of links, and you read them and decide. The human does the work.
- AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity.. You ask a question, the AI pulls in some sources and writes you one answer — an AI OverviewAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index., a ChatGPT reply, a Perplexity answer. The AI reads, you read its summary.
- Agentic search. You give the AI a goal — “compare the best project management tools for a 10-person agency and tell me which to pick” — and it goes off on its own. It plans the work, runs many searches, opens and reads real pages, changes its approach based on what it finds, and finally writes you a report (or even takes an action, like filling a form). You might never see the steps in between.
The cleanest way to remember the relationship: all agentic search is AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity., but not all AI search is agentic. Agentic is the subset that goes off and works on its own for a while.
A concrete example
Same product, two modes:
- ChatGPT, regular answer: you ask something, it replies in a few seconds from what it already “knows” plus maybe a quick web look. One pass.
- ChatGPT Deep Research: you ask the same thing, it says it’ll take a while, and it spends minutes searching, reading, and cross-checking dozens of sources before it writes up a structured answer with citations.
That second mode is agentic search. Google’s Gemini Deep Research and Perplexity’s Deep Research work the same way.
Why an SEO should care
Two reasons:
- These agents read your actual pages, deeply. A traditional crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. skims a lot of pages shallowly. An agent might read 5–10 pages very thoroughly and pull specific claims out of them. Being one of those few pages is the new game.
- You can’t always keep them out. The “research crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.” versions (like
Gemini’s) announce themselves and follow your
robots.txt. But the agentic browser versions — ChatGPT Agent, Google’s Project Mariner — open a real Chrome browser in the cloud and look exactly like a person visiting. There’s no bot name to block, sorobots.txtdoes nothing.
Want the full breakdown — every platform, the user-agentA user agent is the HTTP request header a client (browser, crawler, or bot) sends to identify itself. For crawlers, a short user-agent token — a substring of that string — is what robots.txt rules actually target. table, and what content signals actually help — switch to the Advanced tab.
TL;DR — Agentic searchAgentic search is an AI system that autonomously plans, runs, and synthesizes a multi-step research or browsing task — dozens to hundreds of searches and real page reads — instead of returning one answer to a single query. All agentic search is AI search; not all AI search is agentic. is the third layer of search (traditional → AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. → agentic): a system that runs a multi-step loop — Plan → Search → Read → Evaluate → re-search → synthesize/act — over minutes, firing 20–160+ queries per session and reading full page content, not cached snapshots. The platforms split into two camps for control purposes. Declared crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. (
Gemini-Deep-Research,Claude-User) respectrobots.txt. Agentic browsers (ChatGPT Agent, Project Mariner, Copilot Actions) spin up real Chrome/Edge VMs with no declared bot UA —robots.txtcan’t touch them; only server-side controls or the emerging Web Bot Auth RFC can. The payoff: AI-referred traffic converts ~42% better than non-AI, and agentic traffic is the highest-intent slice of that.
Agentic search is a third layer, not a replacement
Current vendor descriptions support multi-step research and tool use, not a universal replacement for retrieval or ranking. Evidence for this claim Deep-research products can perform multi-step web research and synthesize cited results. Scope: Documented product capabilities, not a universal definition or ranking architecture for agentic search. Confidence: high · Verified: OpenAI: Introducing deep research Pipeline diagrams here are explanatory models unless explicitly documented by a provider. Evidence for this claim The term agentic search covers an evolving set of vendor implementations whose tools and behavior can differ. Scope: Editorial synthesis from documented product behavior; implementations and releases may vary. Confidence: medium · Verified: OpenAI: Introducing deep research
The mental model I keep coming back to:
- Traditional search returns links; the human reads and decides.
- AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. (the GEOGenerative Engine Optimization (GEO) is the practice of optimizing content and brand presence so AI-powered search engines and assistants — Google AI Overviews, ChatGPT, Perplexity — cite, recommend, or mention you when generating answers. Google's position is that it's still SEO./AEOAnswer Engine Optimization (AEO) is the practice of structuring content so engines deliver it as a direct answer — featured snippets, voice assistants, and AI search — rather than just a ranked link. Coined for voice search in 2018 and revived for the LLM era. Google's position is that it's still SEO. target — AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index., ChatGPT answers, Perplexity) returns one answer synthesized from retrieved context.
- Agentic search breaks a goal into sub-tasks, picks its own tools, reads full pages, iterates on what it learns, and ends in a synthesized report or an action (a purchase, a booking, a form fill).
Semrush put the relationship as well as anyone: “All agentic search is AI search. Not all agentic search is the same thing as AI search.” The short version that sticks — all agentic search is AI search; not all AI search is agentic. None of this replaces traditional SEO. Authority, indexability, and structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. are still prerequisites; agentic optimization stacks completeness, machine-readability, and entity clarity on top.
How agentic systems actually work
The defining difference from RAGRAG is the retrieve-then-generate pattern behind AI search: the system retrieves relevant passages from an external index at query time, injects them into the model's context, and generates an answer grounded in those sources — without changing the model's weights. or a single retrieval call is that retrieval happens mid-reasoning, on a loop:
An agent begins by breaking a goal into sub-tasks. It formulates one or more searches, fetches and reads live sources, and evaluates gaps, conflicts, and evidence. When evidence is missing, it revises the query and loops back to search. Once the task is sufficiently supported, it synthesizes a report or takes an allowed action. This is an explanatory model; products vary and do not all expose or implement every step identically.
Firecrawl describes it as a method where the agent “decides when to search, formulates the query from context, reads actual page content from each result, and evaluates what it foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. before deciding whether to synthesize or search again.” The practical consequences:
- Real-time page reads, not pre-indexedStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. summaries. Agents fetch and read full live content. Freshness and live availability matter more than for traditional search; a cached snapshot isn’t what they’re working from.
- Iterative query reformulation. If results are stale or off-target, the agent rewrites the query (e.g. tacking “2026” onto a query when it gets deprecated docs) and goes again.
- Tool chaining. A single task can combine web search, code execution, file reads, and form submission.
This is a deterministic coverage simulation, not a reproduction of any platform's private agent or ranking system.
Generate a query fan-out and compare its branches with a page or passage using my free Query Fan-Out Simulator Free
- Enter the goal-like query an agent would need to research and the page or passage you want to test.
- Review uncovered and partially covered sub-queries instead of treating one broad topic match as complete coverage.
- Confirm the answer shape and evidence manually; this run uses lexical BM25 and cosine components and does not evaluate embeddings.
The completed Query Fan-Out Simulator result contains five sub-queries for technical SEO audit. Two branches show 100 percent term coverage, two show 60 percent, and one shows 75 percent. Each result says to confirm the answer shape and evidence. The result also states that coverage uses lexical BM25 and cosine components only and that embeddings are not evaluated in this run.
How many searches per session
This is not one query. Volume by platform:
| Product | Searches per session | Duration |
|---|---|---|
| Perplexity Deep Research | 20–50 queries | 2–4 minutes |
| Gemini Deep Research | ~80 (standard); ~160 (Max) | up to 60 min (most < 20) |
| OpenAI Deep Research | unspecified; “tens of minutes” | 10–30+ minutes |
| Project Mariner | task-based browsing (up to 10 parallel tasks) | per task |
| ChatGPT Agent | up to ~10 parallel tasks | per task |
Google’s SAGE research found AI agents take an average of 4.9 steps per query — searching, comparing, and evaluating across multiple sources before delivering a result.
The major agentic platforms
OpenAI — Deep Research + ChatGPT Agent. Deep Research (o3-deep-research /
o4-mini-deep-research) is built to “find, analyze, and synthesize hundreds of
sources to create a comprehensive report at the level of a research analyst,” and
OpenAI warns the requests “can take a long time.” ChatGPT Agent is the successor
to Operator — a cloud-hosted Chromium instance that browses the real web and can run
~10 parallel tasks. Its retrieval bots (OAI-SearchBot, ChatGPT-User) are
declared; the agentic browser is not.
Google — Gemini Deep Research + Project Mariner. Gemini Deep Research runs an
explicit “Plan → Search → Read → Iterate → Output” loop — Google describes it
“continuously refin[ing] its analysis, browsing the web by searching and finding
interesting pieces of information then starting a new search based on what it’s
learned, repeating this process multiple times before generating a comprehensive
report.” It has a declared Gemini-Deep-Research user agentA user agent is the HTTP request header a client (browser, crawler, or bot) sends to identify itself. For crawlers, a short user-agent token — a substring of that string — is what robots.txt rules actually target. and respects
robots.txt. Project Mariner is the browser-automation agent (up to 10 parallel
tasks) — it runs a standard Chrome fingerprint from cloud VMs with no declared
bot UA.
Perplexity — Deep Research / Sonar. Perplexity “performs dozens of searches,
reads hundreds of sources, and reasons through the material to autonomously deliver
a comprehensive report” — 20–50 queries in 2–4 minutes. Its declared PerplexityBot
says it respects robots.txt. In August 2025,
Cloudflare reported
observing undeclared crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. it attributed to Perplexity reaching content despite
no-crawl directives; Perplexity disputed Cloudflare’s account.
Separately, Perplexity’s docs say Perplexity-User generally ignores robots.txt
because it represents a user-requested fetch.
Anthropic — Claude web search + computer use. Claude-User (real-time
user-initiated fetches) respects robots.txt. When Claude runs in computer-use
mode it drives a standard browser and emits no declared bot UA — the same agentic-browser
pattern as Mariner and ChatGPT Agent.
Microsoft — Copilot + Copilot Actions. Copilot’s underlying web retrieval uses
BingbotBingbot is Microsoft Bing's primary web crawler — the bot that discovers, fetches, and renders pages to build the Bing index. That index also powers Yahoo, DuckDuckGo, Ecosia, and Microsoft Copilot, so Bingbot's reach is far wider than Bing's own search-market share., which respects robots.txt. Copilot Actions (agentic browsing) uses a
standard Edge/Chromium UA with no declared bot signal — indistinguishable from
human browsing at the UA layer.
User agents and the robots.txt gap
This is the part that trips people up, so here’s the whole landscape in one table:
| Bot / agent | User agent signal | robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. | Type |
|---|---|---|---|
| GPTBot | …compatible; GPTBot/1.3 | Respects | Training |
| OAI-SearchBot | …compatible; OAI-SearchBot/1.3 | Respects | Search index |
| ChatGPT-User | …compatible; ChatGPT-User/1.0 | Unclear since Dec 2025 | User-triggered |
| ChatGPT Agent | Standard Chrome UA + Signature-Agent header | No signal | Agentic browser |
| ClaudeBot | …compatible; ClaudeBot/1.0 | Respects | Training |
| Claude-SearchBot | Claude-SearchBot | Respects | Search index |
| Claude-User | …compatible; Claude-User/1.0 | Respects | User-triggered |
| PerplexityBot | …compatible; PerplexityBot/1.0 | Says yes; Cloudflare disputed compliance | Search index |
| Perplexity-User | …compatible; Perplexity-User/1.0 | Generally ignores | User-triggered |
| Gemini-Deep-Research | …compatible; Gemini-Deep-Research; +https://gemini.google/… | Respects | Deep-research agent |
| Google-Agent | Standard Chrome UA | Ignores | User-triggered |
| Bingbot | …compatible; bingbot/2.0 | Respects | Search / Copilot index |
| Copilot Actions | Standard Edge/Chromium UA | No signal | Agentic browser |
| Project Mariner | Standard Chrome UA (cloud VM) | No signal | Agentic browser |
The categories that matter: training crawlers and search-index bots respect
robots.txt; user-triggered fetchers vary (OpenAI quietly muddied
ChatGPT-User’s position in December 2025); and agentic browsers ignore it by
design.
The key gotcha: agentic browsers are unblockable via robots.txt
robots.txt is a directive for crawlers. Agentic browsers — ChatGPT Agent,
Project Mariner, Copilot Actions, computer-use Claude — spin up real Chrome or Edge
instances inside cloud VMs. They present as standard browsers, declare no bot
identity, and are treated as browser proxies rather than crawlers. So:
- You can block training (
GPTBot,ClaudeBot) and search-index bots (OAI-SearchBot,Claude-SearchBot) inrobots.txt. - You cannot block the agentic browsers there. They simply don’t read it.
The only things that work against the agentic-browser layer are server-side
controls (IP-range rules, rate limiting, bot-management at the edge) and the
emerging Web Bot Auth RFC — cryptographic agent verification. ChatGPT Agent is
already an early signal here: it sends an ed25519-signed Signature-Agent: "https://chatgpt.com" HTTP Message Signature, which is the only reliable
server-side identifier for it (in plain analytics it shows up as “direct/(none)” on
a direct visit, or “Bing/organic” when it found you via Bing search).
One corollary worth saying out loud: blocking GPTBot does not stop ChatGPT from
reading your pages in real time. GPTBot is the training crawler. Deep Research
and ChatGPT Agent reach your content through entirely different mechanisms.
What content signals matter for agentic AI
Agents read full pages and evaluate them for completeness, citeability, consistency, and machine-readability. From the research, what actually helps:
- Comprehensive, single-resource pages. Agents go deep but narrow — they may read 5–10 pages thoroughly rather than skim hundreds. A page that fully answers a topic (and the sub-questions it raises) is far more useful than five thin ones. ALM Corp’s framing: “Structure gives agents stable extraction points — a page with clear headings, labeled sections and consistent formatting is easier to quote accurately.”
- Well-cited, authoritative, machine-readable content. Agents cross-reference against third-party sources and knowledge bases, so external citations help you clear their trust threshold. This lines up with what I found for AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index.: mentions on heavily-linked pages correlate at ~0.70 (Spearman) with AI visibility — the strongest single signal in that study, and a pattern I’d expect agentic retrieval to lean on too.
- Structured data. JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. isn’t required for citation, but it improves extraction accuracy — GPT-4 went from 16% to 54% correct responses when content carried structured data in one cited test.
- Speed. BrightEdge data shows ChatGPT agents abandon slow-loading sites immediately — they have less patience than GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer., which extends Core Web Vitals discipline into the agentic context. Content also needs to be reachable without JavaScript execution.
- Scope and tradeoff clarity. Pages that state who a solution is for, who it’s not for, and the tradeoffs are more useful to an agent building a shortlist than content trying to appeal to everyone.
Why this is worth the effort: conversion quality
The traffic numbers are small but the quality is not. As of March 2026, AI-referred traffic converts about 42% better than non-AI traffic, and AI search traffic converts at roughly 14.2% vs. Google’s 2.8%. Agentic traffic is a subset of AI traffic — and almost certainly the highest-intent slice, because an agent only lands a user on your page after evaluating it against alternatives on the user’s behalf. Being one of the handful of pages an agent reads, and being the one it recommends, is disproportionately valuable.
How to think about it: SEO → GEO → ASO
Treat agentic optimization as a third layer stacked on the first two, not a
replacement. The foundations don’t change — authority, indexability, structured
data. What’s new on top: comprehensive single-topic resources, ruthless
machine-readability, entity clarity, external citations, and the recognition that
robots.txt is no longer a complete access-control story. See the Checklists tab
for the concrete pass.
AI summary
A condensed take on the Advanced version:
- Agentic searchAgentic search is an AI system that autonomously plans, runs, and synthesizes a multi-step research or browsing task — dozens to hundreds of searches and real page reads — instead of returning one answer to a single query. All agentic search is AI search; not all AI search is agentic. is search’s third layer: traditional (links) → AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. (one answer) → agentic (autonomous multi-step research/action). “All agentic search is AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity.. Not all AI search is agentic.”
- It runs a loop, not a query: Plan → Search → Read → Evaluate → re-search → synthesize/act, over minutes. Gemini Deep Research fires ~80–160 queries/session; Perplexity 20–50; Google’s SAGE foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. ~4.9 steps per query.
- Two control camps. Declared crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. (
Gemini-Deep-Research,Claude-User) respectrobots.txt. Agentic browsers (ChatGPT Agent, Project Mariner, Copilot Actions) run real Chrome/Edge VMs with no declared bot UA —robots.txtcan’t block them. - The gotcha: only server-side controls or the emerging Web Bot Auth RFC stop
agentic browsers. ChatGPT Agent’s ed25519
Signature-Agentheader is the only reliable server-side ID. BlockingGPTBot≠ stopping ChatGPT from reading you. - Cloudflare and Perplexity dispute what happened. Cloudflare reported that undeclared crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. it attributed to Perplexity bypassed no-crawl directives; Perplexity denied the allegation. Treat that episode as a reported dispute, not a settled finding.
- What helps: comprehensive single-topic pages, external citations (~0.70 Spearman with AI visibilityLLM visibility (or AI visibility) is the aggregate measure of how often and how prominently a brand or page shows up in AI-generated answers — across AI Overviews, ChatGPT, Perplexity, Copilot, and Gemini. It's the AI-search analog of organic visibility, but it's driven by different signals.), structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding., speed (agents abandon slow sites), scope clarity.
- Why bother: AI-referred traffic converts ~42% better; AI search 14.2% vs. Google’s 2.8%. Agentic is the highest-intent subset.
Official documentation
Primary-source documentation from the platform makers.
OpenAI
- Deep Research guide — the
o3-deep-research/o4-mini-deep-researchmodels, background execution, and tool access. - OpenAI bots & crawlers — full user-agentA user agent is the HTTP request header a client (browser, crawler, or bot) sends to identify itself. For crawlers, a short user-agent token — a substring of that string — is what robots.txt rules actually target. strings for GPTBot, OAI-SearchBot, and ChatGPT-User.
- Introducing ChatGPT Agent — the agentic browser successor to Operator.
- Gemini Deep Research API docs — query counts, duration limits, and background execution.
- Try Deep Research in Gemini — Google’s plain-language description of the research loop.
Anthropic
- Claude computer-use tool — how Claude drives a browser/desktop in agentic mode.
Perplexity
- Sonar Deep Research — the API model behind Deep Research.
- Perplexity crawlers — PerplexityBot vs. Perplexity-User and their stated robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. behavior.
Microsoft
- Bing Webmaster Tools — BingbotBingbot is Microsoft Bing's primary web crawler — the bot that discovers, fetches, and renders pages to build the Bing index. That index also powers Yahoo, DuckDuckGo, Ecosia, and Microsoft Copilot, so Bingbot's reach is far wider than Bing's own search-market share. underpins Copilot’s web retrieval; crawl control lives here.
Quotes from the source
On-the-record statements from the platform makers and the researchers documenting their behavior.
Google — the Gemini Deep Research loop
- “Over the course of a few minutes, Gemini continuously refines its analysis, browsing the web by searching and finding interesting pieces of information then starting a new search based on what it’s learned, repeating this process multiple times before generating a comprehensive report.” — Google Blog. Source
OpenAI — what Deep Research is for
- Deep Research models “find, analyze, and synthesize hundreds of sources to create
a comprehensive report at the level of a research analyst.”
[unverified]— paraphrase of OpenAI’s Deep Research positioning; confirm against the live doc. Source
Perplexity — how Deep Research works
- “Perplexity performs dozens of searches, reads hundreds of sources, and reasons through the material to autonomously deliver a comprehensive report.” — Perplexity Blog. Source
Cloudflare — Perplexity’s stealth crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.
Cloudflare’s allegation is quoted below; Perplexity disputed the account.- “Both their declared and undeclared crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. were attempting to access the content for scraping contrary to the web crawlingCrawling is how search engines use automated bots (like Googlebot and Bingbot) to discover URLs and download pages. A page has to be crawlable to be indexed, but crawling on its own isn't a ranking factor. norms as outlined in RFC 9309.” — Cloudflare. Source
Firecrawl — defining agentic searchAgentic search is an AI system that autonomously plans, runs, and synthesizes a multi-step research or browsing task — dozens to hundreds of searches and real page reads — instead of returning one answer to a single query. All agentic search is AI search; not all AI search is agentic.
- Agentic search is a method where the agent “decides when to search, formulates the query from context, reads actual page content from each result, and evaluates what it foundA 302 (\"Found\") is a temporary redirect: it forwards users to a new URL while telling search engines the original URL should stay in the index. It's a weak canonicalization signal, not the zero-equity dead end of SEO folklore. before deciding whether to synthesize or search again.” — Firecrawl Blog. Source
ALM Corp — what agents reward
- “Structure gives agents stable extraction points — a page with clear headings, labeled sections and consistent formatting is easier to quote accurately.” — ALM Corp. Source
Semrush — the layer distinction
- “All agentic search is AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity.. Not all AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. is agentic.”
[unverified]— paraphrase of Semrush’s framing; confirm exact wording against the live post. Source
[unverified] are paraphrased and should be
confirmed against the live source before being treated as verbatim. Agentic-search optimization checklist
A pass to make sure agents can find, read, trust, and cite you — and that you can see them when they do:
- Comprehensive single-topic pages. Each important topic has one page that fully answers it (and the sub-questions it raises), rather than several thin pages. Agents read deep, not wide.
- Cited sources and external authority. Claims are attributable; you’re mentioned on heavily-linked, high-authority third-party pages (the strongest AI-visibility signal I’ve measured).
- Structured dataStructured data is a standardized way of labeling page content (using the schema.org vocabulary in JSON-LD, Microdata, or RDFa) so search engines can understand its meaning. It's not a direct ranking factor — its value is rich results and entity understanding. in place. JSON-LDJSON-LD (JavaScript Object Notation for Linked Data) is a script-based structured data format, typically paired with the schema.org vocabulary to describe page content for search engines and AI systems. Google recommends it over Microdata and RDFa because it's the easiest format to implement and maintain at scale — but all three work, and structured data isn't a ranking signal. (Organization, FAQPage, Product, etc.) so agents extract facts accurately without guessing.
- Scope and tradeoff clarity. Pages state who a solution is for, who it’s not for, and the tradeoffs — easier for an agent building a shortlist.
- Reachable without JavaScript and fast. Core content in raw HTML; pages load quickly (agents abandon slow sites faster than GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer. does).
- Server-side access controls, not just
robots.txt. If you need to gate access, use IP-range rules / edge botA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index.-management — agentic browsers ignorerobots.txt. -
robots.txtstill set correctly for the declared bots. Decide training (GPTBot,ClaudeBot) vs. retrieval (OAI-SearchBot,Claude-SearchBot) vs. deep-research (Gemini-Deep-Research) per your strategy. - Monitor agentic crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. in server logsLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened.. Watch for the declared deep-research
UAs and the
Signature-Agentheader (ChatGPT Agent’s only reliable ID); remember direct agent visits can show as “direct/(none)” in analytics.
Agentic AI user-agent reference — cheat sheet
| Platform / agent | User-agentA user agent is the HTTP request header a client (browser, crawler, or bot) sends to identify itself. For crawlers, a short user-agent token — a substring of that string — is what robots.txt rules actually target. signal | Category | robots.txtA plain-text file at the root of a host that tells crawlers which URLs they may and may not request. It controls crawling, not indexing — a blocked URL can still be indexed if it's linked from elsewhere. | Notes |
|---|---|---|---|---|
| GPTBot | …compatible; GPTBot/1.3 | Training | Respects | Blocking ≠ stopping ChatGPT live reads |
| OAI-SearchBot | …compatible; OAI-SearchBot/1.3 | Search indexStoring a crawled page in the search index so it can appear in results. Crawled is not the same as indexed — Google selects what to keep, and indexing isn't guaranteed. | Respects | Feeds ChatGPT search |
| ChatGPT-User | …compatible; ChatGPT-User/1.0 | User-triggered | Unclear (Dec 2025 change) | Policy language softened |
| ChatGPT Agent | Chrome UA + Signature-Agent header | Agentic browser | No | ed25519-signed; only reliable server-side ID |
| ClaudeBot | …compatible; ClaudeBot/1.0 | Training | Respects | — |
| Claude-SearchBot | Claude-SearchBot | Search index | Respects | — |
| Claude-User | …compatible; Claude-User/1.0 | User-triggered | Respects | — |
| Claude (computer use) | Standard browser UA | Agentic browser | No | Drives a real browser |
| PerplexityBot | …compatible; PerplexityBot/1.0 | Search index | Says yes; compliance disputed | Cloudflare allegation; Perplexity denial |
| Perplexity-User | …compatible; Perplexity-User/1.0 | User-triggered | Generally ignores | Per Perplexity’s own docs |
| Gemini-Deep-Research | …compatible; Gemini-Deep-Research; +https://gemini.google/… | Deep-research agent | Respects | Declared, blockable |
| Google-Agent | Standard Chrome UA | User-triggered | Ignores | Browser proxy |
| BingbotBingbot is Microsoft Bing's primary web crawler — the bot that discovers, fetches, and renders pages to build the Bing index. That index also powers Yahoo, DuckDuckGo, Ecosia, and Microsoft Copilot, so Bingbot's reach is far wider than Bing's own search-market share. | …compatible; bingbot/2.0 | Search / Copilot index | Respects | Underpins Copilot retrieval |
| Copilot Actions | Standard Edge/Chromium UA | Agentic browser | No | No botA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. signal |
| Project Mariner | Standard Chrome UA (cloud VM) | Agentic browser | No | Up to 10 parallel tasks |
The one-line rule: declared crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. and search bots respect robots.txt;
agentic browsers (real Chrome/Edge VMs) don’t — block those server-side or via
Web Bot Auth.
Agentic-search mistakes
Assuming robots.txt controls every agent
Some agentic systems browse through ordinary browser infrastructure rather than a declared crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. identity. Use robots rules for bots that honor them, but do not present the file as an access-control boundary.
Designing only for one-shot answers
Agents may compare, revisit, filter, and complete multi-step tasks. Keep navigation, forms, states, prices, policies, and entity relationships explicit enough to survive a longer workflow.
Hiding critical facts behind fragile interactions
An agent cannot reliably use information that appears only after an inaccessible gesture, canvas state, or undocumented client action. Preserve semantic HTMLSemantic HTML is the practice of using elements like <main>, <article>, <section>, <nav>, <header>, and <aside> to describe what content is, not just how it looks. It helps search engines and assistive tech identify a page's main content more reliably, but it isn't a ranking factor on its own., stable URLs, clear labels, and server-verifiable outcomes.
The Discover, Understand, Act, Verify framework
- Discover: Important pages and capabilities have stable links or documented machine-readable entry points.
- Understand: Entities, constraints, prices, policies, and next steps are stated clearly in semantic content.
- Act: Forms and workflows have explicit labels, predictable controls, and recoverable error states.
- Verify: A completed action produces a durable confirmation that a user or agent can inspect before continuing.
Use the framework to review a full task journey. A page can be easy to retrieve yet still fail when an agent tries to act or confirm the result.
Patrick's relevant free tools
- Agent Readiness Checker — Check public MCP, agent-card, llms.txt, security, and Content-Signals discovery markers.
- AI Content Brief Generator — Assemble an exportable brief while preserving which research inputs are observed, heuristic, or not evaluated.
Tools for agent-accessible search experiences
- SEO MCP Server — inspect a bounded, read-only example of exposing published SEO content and verification capabilities directly to agents.
- robots.txt Tester — verify declared crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. rules and identify the winning directive, while remembering that ordinary browser-based agents may not present a controllable crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. identity.
- Browser developer tools — inspect semantic names, focus order, network calls, error states, and durable completion signals across the task.
- Server logsLog file analysis is reading a web server's raw access logs to see exactly which URLs search engine crawlers actually requested, when, how often, and what status code they got. Unlike crawl tools or Search Console, logs are the unsampled, ground-truth record of what really happened. — distinguish documented bots, ordinary browsers, and unknown clients without claiming that every request identifies an agent or its purpose.
Test yourself: Agentic search
Resources worth your time
Official platform docs
- OpenAI Deep Research and OpenAI bots.
- Google Gemini Deep Research API and the launch blog post.
- Anthropic Claude computer use.
- Perplexity Sonar Deep Research and Perplexity crawlers.
My related writing
- What We Actually Know About Optimizing for LLM Search — the correlation study behind the ~0.70 Spearman finding for mentions on high-authority pages.
- Meet the New Web Crawlers: AI Bots Are Closing in on Search Engine Bots — the shifting crawlerA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. landscape.
From others
- Cloudflare: Perplexity is using stealth, undeclared crawlers — Cloudflare’s allegation and technical account; Perplexity denied it.
- Simon Willison: ChatGPT agent’s user-agent — the
Signature-Agent/ ed25519 detail. - Firecrawl: What is agentic search — the technical anatomy of the loop.
- Semrush: Agentic search and ALM Corp: The Agentic Web Explained — strategy framing.
- No Hacks: The AI User-Agent Landscape in 2026 and SEJ’s AI crawler user-agent list.
- BrightEdge: Agentic AI Activity Doubles — Adapt Your SEO Strategy Now — data on ChatGPT agent activity doubling between July–August 2025 and agents abandoning slow-loading sites faster than GooglebotGooglebot is Google's web crawler — the software that fetches pages so Google can index and rank them. It comes in two variants, Googlebot Smartphone (primary, under mobile-first indexing) and Googlebot Desktop, and runs an evergreen Chromium renderer..
- Seer Interactive: How Will You Know When OpenAI’s Operator Agent Hits Your Website — practical guide to identifying and measuring agentic browser visits in analytics.
- Cloudflare: From Googlebot to GPTBot — Who’s Crawling Your Site in 2025 — Cloudflare’s crawl-volume data across declared AI botsAI crawlers are bots from AI companies that fetch web pages to train language models, build AI-search indexes, or answer live user questions. They come in three categories, each with its own user-agent tokens and its own robots.txt controls. vs. traditional search crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index..
- Navoto: Agentic Search Optimization — Guide to AI Search Visibility in 2026 — practical ASO framework covering the SEO → GEOGenerative Engine Optimization (GEO) is the practice of optimizing content and brand presence so AI-powered search engines and assistants — Google AI Overviews, ChatGPT, Perplexity — cite, recommend, or mention you when generating answers. Google's position is that it's still SEO. → ASO layering and content signals.
Stats worth citing
- AI-referred traffic converts ~42% better than non-AI traffic (March 2026) — the headline reason agentic visibility is worth the work. Source
- AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. converts at ~14.2% vs. Google’s ~2.8% — agentic traffic is the highest-intent subset of that AI traffic. Source
- Gemini Deep Research: ~80 queries per session (standard), ~160 (Max) — the scale of a single agentic research run. Source
- Perplexity Deep Research: 20–50 queries per session, 2–4 minutes — a lighter but still multi-step run. Source
- ~0.70 Spearman correlation between mentions on heavily-linked pages and AI visibility (Google AI OverviewsAI Overviews are the AI-generated summary box Google shows above or within its regular search results, written by Gemini models from pages retrieved out of Google's normal Search index. It's a Search feature, not a separate platform or index.) — my finding, and the likely shape of agentic trust signals too. Source
- Agentic AI traffic grew 1,300% in the first 8 months of 2025, with ChatGPT agent activity alone doubling between July and August 2025 — the fastest-growing slice of AI-referred traffic. Source
Agentic Search
Agentic search is an AI system that autonomously plans, runs, and synthesizes a multi-step research or browsing task — dozens to hundreds of searches and real page reads — instead of returning one answer to a single query. All agentic search is AI search; not all AI search is agentic.
Related: AI Search, AI Crawlers
Agentic Search
Agentic search is an AI system that autonomously plans, executes, and synthesizes a multi-step research or browsing task on a user’s behalf — without the user directing each individual step. It sits a layer above both traditional search and basic AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity.: a traditional engine returns a ranked list of links for a human to read; an AI searchAI search uses large language models and retrieval-augmented generation (RAG) to synthesize an answer from multiple sources rather than returning a ranked list of links. Examples include Google AI Overviews, ChatGPT Search, and Perplexity. answer generates one response from retrieved context; an agentic system breaks a goal into sub-tasks, decides which tools to use (web search, page reads, code execution, form-filling), iterates based on what it learns, and either delivers a synthesized report or takes an action.
The shorthand from Semrush captures it: “All agentic search is AI search. Not all AI search is agentic.” The defining trait is the loop — Plan → Search → Read → Evaluate → (re-search) → Synthesize/Act — running over minutes. Gemini Deep Research runs roughly 80–160 search queries per session; Perplexity Deep Research runs 20–50; OpenAI Deep Research can run for tens of minutes.
The technical gotcha that matters for site owners: the agentic browsers in this category (ChatGPT Agent, Google’s Project Mariner, Microsoft Copilot Actions) spin up real Chrome or Edge instances in cloud VMs and browse as standard browsers with no declared botA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index. user agentA user agent is the HTTP request header a client (browser, crawler, or bot) sends to identify itself. For crawlers, a short user-agent token — a substring of that string — is what robots.txt rules actually target.. They are treated as browser proxies, not crawlersA crawler — also called a spider or bot — is an automated program that fetches web pages, extracts their links, and queues new URLs to visit. Search engines use crawlers to discover and download content for their index., so robots.txt does not stop them — only server-side controls or the emerging Web Bot Auth verification do. Declared deep-research crawlers like Gemini-Deep-Research and user-fetchers like Claude-User still respect robots.txt; Perplexity’s compliance has been documented as inconsistent.
Related: AI Search, AI Crawlers
Build-time retrieval analysis plus live signals for this exact article. The automatic chunk report includes a deterministic readiness score and is ready without a model download.