How modern web platforms manage AI search crawlers, semantic markdown indexing, and AI answer engine discoverability across the ChatGPT, Perplexity, and Claude era.
Web administrators increasingly distinguish between training crawlers (which scrape content for foundational models) and search crawlers (which cite and link back to publishers). While 41.6% of SaaS companies block GPTBot, only 19.2% block OAI-SearchBot.
| User-Agent / Bot | Allowed (%) | Blocked (%) | Ratio |
|---|---|---|---|
| OAI-SearchBot (ChatGPT Search) | 80.8% | 19.2% | |
| PerplexityBot (Perplexity AI) | 71.2% | 28.8% | |
| ClaudeBot (Anthropic) | 61.8% | 38.2% | |
| GPTBot (OpenAI Training) | 58.4% | 41.6% | |
| Applebot-Extended | 54.2% | 45.8% | |
| Google-Extended (Gemini Training) | 51.4% | 48.6% | |
| Bytespider (ByteDance) | 34.6% | 65.4% |
Originally proposed to provide LLMs with concise, markdown-formatted indices of technical documentation, /llms.txt has grown from an experimental convention into standard AI search infrastructure.
/llms.txt file length: 1,420 tokens (~5,680 characters), ideal for instant prompt-injection into context windows./llms.txt experienced 3.2× higher citation accuracy in Perplexity Pro and ChatGPT Search compared to raw HTML scraping.Permission rates diverge dramatically across software sectors. Developer tools companies welcome AI search crawlers at a 92% rate, whereas Healthcare and Fintech platforms enforce much stricter disallows due to compliance concerns.
| Vertical | Search Bots Allowed | Training Bots Blocked | /llms.txt Adoption |
|---|---|---|---|
| Developer Tools & Infra | 92.4% | 34.2% | 28.4% |
| Fintech & Payments | 74.1% | 52.8% | 6.2% |
| E-Commerce & Retail SaaS | 83.6% | 38.0% | 4.8% |
| Healthcare & Regulated SaaS | 58.2% | 68.4% | 1.2% |
Copy citations in your preferred format for academic papers, technical documentation, or LLM prompting: