Lesson 8 / 25

AI Crawlers and robots.txt

Understand the main AI user agents and decide deliberately what to allow or block.

Different bots, different purposes

AI companies run several crawlers. Examples at the time of writing: OpenAI uses GPTBot (model training), OAI-SearchBot (indexing for ChatGPT search) and ChatGPT-User (fetching on a user's request); Anthropic uses ClaudeBot, Claude-SearchBot and Claude-User; Perplexity uses PerplexityBot and Perplexity-User; Google offers the Google-Extended token to control use of your content for Gemini; CCBot is Common Crawl, whose data many models use. Names and behaviour change, so check each provider's documentation. robots.txt is a voluntary standard that well-behaved crawlers honour. It is a trade-off: blocking can protect content from training, but blocking search-style crawlers can remove you from answers that rely on retrieval. Decide per bot, on purpose.

Be reachable and readable

If crawlers cannot reach or read your pages, nothing else you do will help.

Four checks: allowed, rendered, structured, discoverable.
Figure 3.1 — Allowed, rendered, structured and discoverable.

Testing rules, run

I ran this with Python's urllib.robotparser. GPTBot may fetch the blog but not /private/; ClaudeBot is blocked entirely; other bots like PerplexityBot follow the * rules. Always test your real robots.txt this way after editing.

from urllib import robotparser
robots = """User-agent: GPTBot
Disallow: /private/

User-agent: ClaudeBot
Disallow: /

User-agent: *
Disallow: /admin/
"""
rp = robotparser.RobotFileParser()
rp.parse(robots.splitlines())
for ua, path in (("GPTBot", "/blog/post"), ("GPTBot", "/private/x"),
                 ("ClaudeBot", "/blog/post"), ("PerplexityBot", "/blog/post"),
                 ("PerplexityBot", "/admin/")):
    print(ua, path, rp.can_fetch(ua, "https://example.com" + path))

Output:

GPTBot /blog/post True
GPTBot /private/x False
ClaudeBot /blog/post False
PerplexityBot /blog/post True
PerplexityBot /admin/ False

robots.txt is not security

It asks bots politely; it does not stop a bad actor or protect private data. Anything truly confidential belongs behind authentication, not just a Disallow line.

Quick check: What is a trade-off of blocking search-style AI crawlers?

  • It improves your SEO guaranteed
  • Your site becomes faster for users
  • You may disappear from answers that rely on live retrieval
  • Nothing changes at all
Answer

You may disappear from answers that rely on live retrieval — If a retrieval system cannot fetch your pages, it cannot quote them.