AI Crawlers Explained — GPTBot, ClaudeBot, PerplexityBot, and Which to Allow
A plain-English guide to AI crawlers in 2026—what GPTBot, OAI-SearchBot, ClaudeBot, PerplexityBot, and others do, how to see which ones visit your site, and how to decide what to allow.
Your server logs are full of bots with names like GPTBot, ClaudeBot, and PerplexityBot. Some collect training data. Some power AI search results that can send you visitors. Some only fetch a page when a user asks an assistant to read it. Blocking all of them—or allowing all of them—without understanding the difference can cost you visibility or control.
Table of Contents
- Three Types of AI Crawlers
- Common AI User Agents
- How to See Which AI Crawlers Visit You
- Allow or Block? A Decision Guide
- robots.txt Examples
- Watch Out for Your CDN
- Key Takeaways
Three Types of AI Crawlers
| Type | Purpose | Effect of blocking |
|---|---|---|
| Training crawlers | Collect content to train future models | Your content isn't used for training; future model knowledge of you may be thinner |
| Search / retrieval crawlers | Index pages for AI search answers and citations | You may not be cited or linked in AI answers |
| User-triggered fetchers | Fetch a page when a user asks the assistant to | Assistants can't read your page on request |
Common AI User Agents
| User agent | Operator | Type |
|---|---|---|
| GPTBot | OpenAI | Training |
| OAI-SearchBot | OpenAI | Search |
| ChatGPT-User | OpenAI | User-triggered |
| ClaudeBot | Anthropic | Training |
| Claude-SearchBot | Anthropic | Search |
| Claude-User | Anthropic | User-triggered |
| PerplexityBot | Perplexity | Search |
| Perplexity-User | Perplexity | User-triggered |
| Google-Extended | robots.txt token controlling AI training use (not a separate crawler) | |
| Applebot-Extended | Apple | robots.txt token for AI training use |
| CCBot | Common Crawl | Open dataset used by many models |
Vendors add and rename agents—check each operator's documentation for the current list.
How to See Which AI Crawlers Visit You
Crawlers don't run JavaScript the way browsers do, so most client-side analytics never see them. Options:
- Server or CDN logs filtered by user agent.
- A tracking pixel or server-side beacon that records bot visits.
- DeepSync [AI visibility](/docs/ai-visibility) (Business and Enterprise) shows which AI crawlers read your site and how often, next to the AI referral traffic you receive.
Allow or Block? A Decision Guide
A common middle ground
Allow search and user-triggered agents (so you can be cited and linked), and decide on training crawlers based on your content strategy.
- Marketing sites, SaaS, e-commerce: usually allow search and user-triggered agents—you want visibility. Many also allow training to be well represented in future models.
- Publishers monetizing content: often block training crawlers, allow search agents selectively.
- Private or paywalled content: block, and enforce at the server—robots.txt is a request, not access control.
robots.txt Examples
Allow AI search and user fetchers, block training:
User-agent: GPTBot
Disallow: /
User-agent: ClaudeBot
Disallow: /
User-agent: OAI-SearchBot
Allow: /
User-agent: ChatGPT-User
Allow: /
User-agent: PerplexityBot
Allow: /Watch Out for Your CDN
Some CDNs and security products block AI crawlers by default or inject a managed robots.txt. Your own file may say "allow" while the edge says "block." After changing settings, verify that crawler hits actually appear in your logs or analytics.
Key Takeaways
- Training, search, and user-triggered crawlers serve different purposes.
- Blocking search agents can remove you from AI citations.
- Client-side analytics miss crawlers—use logs or a dedicated tool.
- Check your CDN settings, not just robots.txt.
Conclusion
AI crawlers are the new search spiders. Decide deliberately which ones you let in—and verify that your infrastructure agrees.
Frequently Asked Questions
Related articles
Is ChatGPT Recommending Your Brand? How to Check and Improve AI Visibility (GEO Guide)
11 min readHow to Track ChatGPT, Perplexity, and AI Assistant Traffic to Your Website (2026)
10 min readWhy Is My Direct Traffic So High? 9 Causes of "Dark Traffic" and How to Fix Attribution
9 min readStay in the loop
Get the latest insights on product analytics and user behavior delivered to your inbox.



