Skip to content

AI Crawler User Agents

Which AI bots visit your site, and how to control them.

About AI crawler user agents

AI companies reach websites in three different ways, and each needs a different decision. Training crawlers such as GPTBot, ClaudeBot and CCBot collect pages that may be used to train models. Search crawlers such as OAI-SearchBot, Claude-SearchBot and PerplexityBot index pages so AI assistants can cite and link to them, which can send you visitors. User-triggered fetchers such as ChatGPT-User and Perplexity-User load a single page because a person asked about it, and several operators say robots.txt may not apply to them.

Google-Extended and Applebot-Extended are not crawlers at all: they are robots.txt tokens that control whether content already crawled by Googlebot or Applebot may be used for AI training. Blocking them doesn't remove you from Google Search or Apple's search features. Every entry below comes from the operator's own published documentation, checked in October 2026. Operators add and rename tokens often, so check their pages before relying on a rule, and remember that robots.txt is a request that well-behaved bots honor, not an access control.

18 entries

Documented AI user-agent tokens

TokenOperatorTypeWhat it doesHonors robots.txt?
GPTBotOpenAITrainingCrawls content that may be used to train OpenAI's generative AI foundation models.Yes. Disallowing it signals your content shouldn't be used for training.
OAI-SearchBotOpenAISearchFinds and surfaces websites in ChatGPT's search features. Not used for training.Yes. OpenAI recommends allowing it if you want to appear in ChatGPT search results.
ChatGPT-UserOpenAIUser requestVisits a page for actions a ChatGPT user starts, such as asking about a link.May not apply: OpenAI says robots.txt rules may not apply to user-initiated actions.
OAI-AdsBotOpenAIOtherChecks the safety of landing pages submitted as ads on ChatGPT.Only visits pages submitted as ads; OpenAI doesn't describe robots.txt control.
ClaudeBotAnthropicTrainingCollects public web content that may be used for model development.Yes. Anthropic also honors the non-standard Crawl-delay.
Claude-SearchBotAnthropicSearchIndexes content to improve the quality of search results in Claude.Yes.
Claude-UserAnthropicUser requestFetches pages when a Claude user asks a question that needs them.Yes, according to Anthropic.
Google-ExtendedGoogleTraining controlA control token, not a separate crawler: decides whether content Google crawls may be used to train Gemini models and for grounding.Yes. Doesn't affect inclusion or ranking in Google Search.
Applebot-ExtendedAppleTraining controlA control token, not a separate crawler: decides whether pages Applebot crawls may train Apple's foundation models.Yes. Pages can still appear in Apple's search features through Applebot.
PerplexityBotPerplexitySearchSurfaces and links websites in Perplexity's search results. Perplexity says it isn't used to train foundation models.Yes.
Perplexity-UserPerplexityUser requestVisits a page when a Perplexity user's question needs it.Generally ignored: Perplexity says this fetcher generally ignores robots.txt.
Meta-ExternalAgentMetaTrainingCrawls for training AI models and improving products by indexing content directly.Yes.
Meta-ExternalFetcherMetaUser requestFetches individual links at a user's request in Meta's AI products.May bypass robots.txt, according to Meta.
Meta-WebIndexerMetaSearchAnalyzes content to improve the quality of Meta AI search results.Yes.
AmazonbotAmazonTrainingImproves Amazon's products and services; content may be used to train Amazon AI models.Yes (no Crawl-delay support).
Amzn-SearchBotAmazonSearchImproves search experiences in Amazon products such as Alexa. Not used to train generative AI models.Yes. Without its own group, it follows the rules you give other search bots.
Amzn-UserAmazonUser requestFetches live information when a user's request needs it.May not follow all robots.txt rules, according to Amazon.
CCBotCommon CrawlTrainingBuilds Common Crawl's open web archive, which many AI training datasets are built from.Yes.

Good to know

Block training but stay in AI search

To keep content out of training while still appearing in AI search answers, disallow the training tokens (GPTBot, ClaudeBot, Google-Extended, Applebot-Extended, CCBot, Meta-ExternalAgent) and leave the search crawlers (OAI-SearchBot, Claude-SearchBot, PerplexityBot) allowed. Our robots.txt generator can add these groups for you.

Tokens not listed here

Bytespider (ByteDance) appears in many server logs, but ByteDance publishes no documentation we could verify, so it isn't in the table; block it at your server or CDN if you need to. Older Anthropic tokens such as anthropic-ai and claude-web don't appear in Anthropic's current documentation.

Matching rules

Under RFC 9309, crawlers match their product token case-insensitively and obey the most specific group that names them, ignoring the * group. So a GPTBot group with only Allow: / overrides a Disallow: / under User-agent: *.

More references