Lesson 14 / 34

Indexing Controls: noindex and X-Robots-Tag

Control what appears in search results with meta robots, noindex and X-Robots-Tag, and avoid blocking pages you want removed.

noindex vs Disallow

Disallow in robots.txt stops crawling, but a blocked URL can still be indexed if other sites link to it. To keep a page out of results, let it be crawled and serve noindex. If robots.txt blocks the page, Google never sees the noindex.

Meta robots and X-Robots-Tag

Use a meta tag for HTML pages and the X-Robots-Tag HTTP header for PDFs and other non-HTML files.

<!-- In <head> of the page -->
<meta name="robots" content="noindex, follow">

# HTTP response header (e.g. for PDFs)
X-Robots-Tag: noindex

Use it on thin or private pages

Apply noindex to internal search results, thank-you pages and staging copies. Never noindex pages you want to rank, and re-check after each deployment.