Lesson 10 / 32
Robots.txt
Control crawler access with robots.txt directives.
Robots.txt basics
robots.txt is a file in your root directory (/robots.txt) that tells crawlers which pages to crawl and which to avoid. It's a courtesy; crawlers can ignore it, but most respect it.
Robots.txt syntax
Use User-Agent to target specific crawlers, Disallow to block paths, and Allow to make exceptions. Always allow crawlers to access CSS, JS, and image files.
# All bots
User-agent: *
Allow: /
Disallow: /admin/
Disallow: /private/
Disallow: /temp/
# Block a specific bot
User-agent: BadBot
Disallow: /
# Google bot
User-agent: Googlebot
Allow: /admin/
# Always allow these paths
Allow: /assets/
Allow: /images/
# Sitemap location
Sitemap: https://example.com/sitemap.xml