AI Crawler User Agents
Which AI bots visit your site, and how to control them.
About AI crawler user agents
AI companies reach websites in three different ways, and each needs a different decision. Training crawlers such as GPTBot, ClaudeBot and CCBot collect pages that may be used to train models. Search crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot index pages so AI assistants can cite and link to them, which can send you visitors. User-triggered fetchers such as ChatGPT-User and Perplexity-User load a single page because a person asked about it, and several operators say robots.txt may not apply to them.
Google-Extended and Applebot-Extended are not crawlers at all: they are robots.txt tokens that control whether content already crawled by Googlebot or Applebot may be used for AI training. Blocking them doesn't remove you from Google Search or Apple's search features. Every entry below comes from the operator's own published documentation, checked in October 2026. Operators add and rename tokens often, so check their pages before relying on a rule, and remember that robots.txt is a request that well-behaved bots honor, not an access control.
18 entries
Documented AI user-agent tokens
| Token | Operator | Type | What it does | Honors robots.txt? |
|---|---|---|---|---|
| GPTBot | OpenAI | Training | Crawls content that may be used to train OpenAI's generative AI foundation models. | Yes. Disallowing it signals your content shouldn't be used for training. |
| OAI-SearchBot | OpenAI | Search | Finds and surfaces websites in ChatGPT's search features. Not used for training. | Yes. OpenAI recommends allowing it if you want to appear in ChatGPT search results. |
| ChatGPT-User | OpenAI | User request | Visits a page for actions a ChatGPT user starts, such as asking about a link. | May not apply: OpenAI says robots.txt rules may not apply to user-initiated actions. |
| OAI-AdsBot | OpenAI | Other | Checks the safety of landing pages submitted as ads on ChatGPT. | Only visits pages submitted as ads; OpenAI doesn't describe robots.txt control. |
| ClaudeBot | Anthropic | Training | Collects public web content that may be used for model development. | Yes. Anthropic also honors the non-standard Crawl-delay. |
| Claude-SearchBot | Anthropic | Search | Indexes content to improve the quality of search results in Claude. | Yes. |
| Claude-User | Anthropic | User request | Fetches pages when a Claude user asks a question that needs them. | Yes, according to Anthropic. |
| Google-Extended | Training control | A control token, not a separate crawler: decides whether content Google crawls may be used to train Gemini models and for grounding. | Yes. Doesn't affect inclusion or ranking in Google Search. | |
| Applebot-Extended | Apple | Training control | A control token, not a separate crawler: decides whether pages Applebot crawls may train Apple's foundation models. | Yes. Pages can still appear in Apple's search features through Applebot. |
| PerplexityBot | Perplexity | Search | Surfaces and links websites in Perplexity's search results. Perplexity says it isn't used to train foundation models. | Yes. |
| Perplexity-User | Perplexity | User request | Visits a page when a Perplexity user's question needs it. | Generally ignored: Perplexity says this fetcher generally ignores robots.txt. |
| Meta-ExternalAgent | Meta | Training | Crawls for training AI models and improving products by indexing content directly. | Yes. |
| Meta-ExternalFetcher | Meta | User request | Fetches individual links at a user's request in Meta's AI products. | May bypass robots.txt, according to Meta. |
| Meta-WebIndexer | Meta | Search | Analyzes content to improve the quality of Meta AI search results. | Yes. |
| Amazonbot | Amazon | Training | Improves Amazon's products and services; content may be used to train Amazon AI models. | Yes (no Crawl-delay support). |
| Amzn-SearchBot | Amazon | Search | Improves search experiences in Amazon products such as Alexa. Not used to train generative AI models. | Yes. Without its own group, it follows the rules you give other search bots. |
| Amzn-User | Amazon | User request | Fetches live information when a user's request needs it. | May not follow all robots.txt rules, according to Amazon. |
| CCBot | Common Crawl | Training | Builds Common Crawl's open web archive, which many AI training datasets are built from. | Yes. |
Good to know
Block training but stay in AI search
To keep content out of training while still appearing in AI search answers, disallow the training tokens (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Meta-ExternalAgent) and leave the search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) allowed. Our robots.txt generator can add these groups for you.
Tokens not listed here
Bytespider (ByteDance) appears in many server logs, but ByteDance publishes no documentation we could verify, so it isn't in the table; block it at your server or CDN if you need to. Older Anthropic tokens such as anthropic-ai and claude-web don't appear in Anthropic's current documentation.
Matching rules
Under RFC 9309, crawlers match their product token case-insensitively and obey the most specific group that names them, ignoring the * group. So a GPTBot group with only Allow: / overrides a Disallow: / under User-agent: *.
More references
- Meta tags cheat sheetWhich meta tags Google uses and which it ignores
- Robots meta directivesnoindex, nosnippet, max-snippet and X-Robots-Tag
- Schema types for rich resultsGoogle's rich result types and their required properties
- HTTP status codes for SEOHow Google treats 2xx, 3xx, 4xx and 5xx responses
- Hreflang codesLanguage and region codes for hreflang, and common mistakes
- SEO glossary71 SEO terms explained in plain English