Lesson 14 / 34
Indexing Controls: noindex and X-Robots-Tag
Control what appears in search results with meta robots, noindex and X-Robots-Tag, and avoid blocking pages you want removed.
noindex vs Disallow
Disallow in robots.txt stops crawling, but a blocked URL can still be indexed if other sites link to it. To keep a page out of results, let it be crawled and serve noindex. If robots.txt blocks the page, Google never sees the noindex.
Meta robots and X-Robots-Tag
Use a meta tag for HTML pages and the X-Robots-Tag HTTP header for PDFs and other non-HTML files.
<!-- In <head> of the page -->
<meta name="robots" content="noindex, follow">
# HTTP response header (e.g. for PDFs)
X-Robots-Tag: noindexUse it on thin or private pages
Apply noindex to internal search results, thank-you pages and staging copies. Never noindex pages you want to rank, and re-check after each deployment.