# AI Crawlers and robots.txt — AI Visibility and LLM Brand Discovery

Source: https://www.geekswithgeeks.com/en/ai-visibility/tech-crawlers-robots

> Understand the main AI user agents and decide deliberately what to allow or block.

## Different bots, different purposes

AI companies run several crawlers. Examples at the time of writing: OpenAI uses **GPTBot** (model training), **OAI-SearchBot** (indexing for ChatGPT search) and **ChatGPT-User** (fetching on a user's request); Anthropic uses **ClaudeBot**, **Claude-SearchBot** and **Claude-User**; Perplexity uses **PerplexityBot** and **Perplexity-User**; Google offers the **Google-Extended** token to control use of your content for Gemini; **CCBot** is Common Crawl, whose data many models use. Names and behaviour change, so check each provider's documentation. **robots.txt** is a voluntary standard that well-behaved crawlers honour. It is a trade-off: blocking can protect content from training, but blocking search-style crawlers can remove you from answers that rely on retrieval. Decide per bot, on purpose.

## Be reachable and readable

If crawlers cannot reach or read your pages, nothing else you do will help.

![Four checks: allowed, rendered, structured, discoverable.](assets/figures/ai-visibility/section-3-map.svg) — Figure 3.1 — Allowed, rendered, structured and discoverable.

## Testing rules, run

I ran this with Python's `urllib.robotparser`. GPTBot may fetch the blog but not `/private/`; ClaudeBot is blocked entirely; other bots like PerplexityBot follow the `*` rules. Always test your real robots.txt this way after editing.

```python
from urllib import robotparser
robots = """User-agent: GPTBot
Disallow: /private/

User-agent: ClaudeBot
Disallow: /

User-agent: *
Disallow: /admin/
"""
rp = robotparser.RobotFileParser()
rp.parse(robots.splitlines())
for ua, path in (("GPTBot", "/blog/post"), ("GPTBot", "/private/x"),
                 ("ClaudeBot", "/blog/post"), ("PerplexityBot", "/blog/post"),
                 ("PerplexityBot", "/admin/")):
    print(ua, path, rp.can_fetch(ua, "https://example.com" + path))
```

Output:

```
GPTBot /blog/post True
GPTBot /private/x False
ClaudeBot /blog/post False
PerplexityBot /blog/post True
PerplexityBot /admin/ False
```

## robots.txt is not security

It asks bots politely; it does not stop a bad actor or protect private data. Anything truly confidential belongs behind authentication, not just a Disallow line.

**Quiz:** What is a trade-off of blocking search-style AI crawlers?

- [ ] It improves your SEO guaranteed
- [ ] Your site becomes faster for users
- [x] You may disappear from answers that rely on live retrieval
- [ ] Nothing changes at all

*Answer:* You may disappear from answers that rely on live retrieval. If a retrieval system cannot fetch your pages, it cannot quote them.
