Lesson 10 / 32

Robots.txt

Control crawler access with robots.txt directives.

Robots.txt basics

robots.txt is a file in your root directory (/robots.txt) that tells crawlers which pages to crawl and which to avoid. It's a courtesy; crawlers can ignore it, but most respect it.

Robots.txt syntax

Use User-Agent to target specific crawlers, Disallow to block paths, and Allow to make exceptions. Always allow crawlers to access CSS, JS, and image files.

# All bots
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /private/
Disallow: /temp/

# Block a specific bot
User-agent: BadBot
Disallow: /

# Google bot
User-agent: Googlebot
Allow: /admin/

# Always allow these paths
Allow: /assets/
Allow: /images/

# Sitemap location
Sitemap: https://example.com/sitemap.xml