Complete technical taxonomy of all 12 major artificial intelligence web crawlers. Understand bot intent, verify reverse DNS signatures, and configure optimal robots.txt policies.
Powers real-time search queries in ChatGPT. Reads content to cite and link back to publishers.
OAI-SearchBot.*\.search\.openai\.comUser-agent: OAI-SearchBot Allow: /
Crawls the web to construct real-time factual citations and search indices for Perplexity Pro.
PerplexityBot.*\.perplexity\.aiUser-agent: PerplexityBot Allow: /
Operated by Anthropic to index web content for Claude web browsing and retrieval.
ClaudeBot.*\.claudebot\.anthropic\.comUser-agent: ClaudeBot Allow: /
Scrapes large-scale web datasets to train future OpenAI foundational models (GPT-4.5, GPT-5).
GPTBot.*\.gptbot\.openai\.comUser-agent: GPTBot Disallow: /
Standalone token allowing publishers to opt out of Gemini and Vertex AI training while remaining indexed in standard Google Search.
Google-Extended.*\.googlebot\.comUser-agent: Google-Extended Disallow: /
High-frequency scraper harvesting content for ByteDance LLMs and TikTok AI applications.
Bytespider.*\.bytespider\.comUser-agent: Bytespider Disallow: /
Permits publishers to opt out of Apple Intelligence generative training while remaining indexed for Siri and Safari search.
Applebot-Extended.*\.applebot\.apple\.comUser-agent: Applebot-Extended Disallow: /
Sent on behalf of an end-user when they paste a specific URL into ChatGPT and request a summary or analysis.
ChatGPT-User.*\.openai\.comUser-agent: ChatGPT-User Allow: /
Crawls web pages to gather training data for Llama open-weights models and Meta AI assistants.
Meta-ExternalAgent.*\.fbsv\.netUser-agent: Meta-ExternalAgent Disallow: /
Indexes enterprise data and public web pages for Cohere Command and Embed models.
cohere-ai.*\.cohere\.aiUser-agent: cohere-ai Disallow: /
Crawls web content for Alexa answer engines, search features, and Amazon Bedrock RAG pipelines.
Amazonbot.*\.amazonbot\.amazon\.comUser-agent: Amazonbot Allow: /
The non-profit Common Crawl foundation crawler used as the foundational corpus for most open-source LLMs.
CCBot.*\.commoncrawl\.orgUser-agent: CCBot Disallow: /