Generative Engine Optimization (GEO)

AI Web Crawlers & Bots Directory

Complete technical taxonomy of all 12 major artificial intelligence web crawlers. Understand bot intent, verify reverse DNS signatures, and configure optimal robots.txt policies.

OpenAI

OAI-SearchBot

Real-Time AI Search

Powers real-time search queries in ChatGPT. Reads content to cite and link back to publishers.

User-Agent Header
OAI-SearchBot
rDNS Verification
.*\.search\.openai\.com
Recommended Policy: ALLOW (Drives high-intent referral clicks)
User-agent: OAI-SearchBot
Allow: /
Perplexity AI

PerplexityBot

Real-Time AI Search

Crawls the web to construct real-time factual citations and search indices for Perplexity Pro.

User-Agent Header
PerplexityBot
rDNS Verification
.*\.perplexity\.ai
Recommended Policy: ALLOW (Primary source of AI conversational search referrals)
User-agent: PerplexityBot
Allow: /
Anthropic

ClaudeBot

Real-Time AI Search & Retrieval

Operated by Anthropic to index web content for Claude web browsing and retrieval.

User-Agent Header
ClaudeBot
rDNS Verification
.*\.claudebot\.anthropic\.com
Recommended Policy: ALLOW (Enables citation inside Claude conversations)
User-agent: ClaudeBot
Allow: /
OpenAI

GPTBot

Foundational Model Pre-Training

Scrapes large-scale web datasets to train future OpenAI foundational models (GPT-4.5, GPT-5).

User-Agent Header
GPTBot
rDNS Verification
.*\.gptbot\.openai\.com
Recommended Policy: DISALLOW (If you wish to prevent model training on proprietary content)
User-agent: GPTBot
Disallow: /
Google

Google-Extended

Gemini Model Pre-Training

Standalone token allowing publishers to opt out of Gemini and Vertex AI training while remaining indexed in standard Google Search.

User-Agent Header
Google-Extended
rDNS Verification
.*\.googlebot\.com
Recommended Policy: DISALLOW (Prevents Gemini training without harming Google ranking)
User-agent: Google-Extended
Disallow: /
ByteDance

Bytespider

ByteDance / TikTok LLM Training

High-frequency scraper harvesting content for ByteDance LLMs and TikTok AI applications.

User-Agent Header
Bytespider
rDNS Verification
.*\.bytespider\.com
Recommended Policy: DISALLOW (Known for aggressive crawl rates and zero attribution)
User-agent: Bytespider
Disallow: /
Apple

Applebot-Extended

Apple Intelligence Model Training

Permits publishers to opt out of Apple Intelligence generative training while remaining indexed for Siri and Safari search.

User-Agent Header
Applebot-Extended
rDNS Verification
.*\.applebot\.apple\.com
Recommended Policy: DISALLOW (Prevents Apple model training without hurting Safari search)
User-agent: Applebot-Extended
Disallow: /
OpenAI

ChatGPT-User

Direct User Browsing Request

Sent on behalf of an end-user when they paste a specific URL into ChatGPT and request a summary or analysis.

User-Agent Header
ChatGPT-User
rDNS Verification
.*\.openai\.com
Recommended Policy: ALLOW (Enables paying ChatGPT users to read your links)
User-agent: ChatGPT-User
Allow: /
Meta

Meta-ExternalAgent

Meta Llama Model Training

Crawls web pages to gather training data for Llama open-weights models and Meta AI assistants.

User-Agent Header
Meta-ExternalAgent
rDNS Verification
.*\.fbsv\.net
Recommended Policy: DISALLOW (If restricting open-weights AI model scraping)
User-agent: Meta-ExternalAgent
Disallow: /
Cohere

cohere-ai

Cohere Command Model Training

Indexes enterprise data and public web pages for Cohere Command and Embed models.

User-Agent Header
cohere-ai
rDNS Verification
.*\.cohere\.ai
Recommended Policy: DISALLOW (Optional opt-out for enterprise LLM scraping)
User-agent: cohere-ai
Disallow: /
Amazon

Amazonbot

Alexa & Bedrock Web Retrieval

Crawls web content for Alexa answer engines, search features, and Amazon Bedrock RAG pipelines.

User-Agent Header
Amazonbot
rDNS Verification
.*\.amazonbot\.amazon\.com
Recommended Policy: ALLOW (Maintains Alexa and AWS answer engine reach)
User-agent: Amazonbot
Allow: /
Common Crawl

CCBot

Open Source AI Web Scraping

The non-profit Common Crawl foundation crawler used as the foundational corpus for most open-source LLMs.

User-Agent Header
CCBot
rDNS Verification
.*\.commoncrawl\.org
Recommended Policy: DISALLOW (If preventing non-attributed bulk corpus inclusion)
User-agent: CCBot
Disallow: /