Lesson 11 / 25
Sitemaps and llms.txt
Help crawlers discover important pages and evaluate the experimental llms.txt idea sensibly.
Make discovery easy
An XML sitemap lists the URLs you want crawled and when they changed; reference it in robots.txt. Keep internal links clear so important pages are reachable in a few clicks. llms.txt is a community proposal for a Markdown file at /llms.txt that gives a short summary of your site with links to your most useful pages, written for language models. As of writing it is an experiment: major providers have not publicly committed to using it, so it is cheap to add but should not replace good pages and sitemaps. Check the current status before relying on it.
Parsing an llms.txt, run
I ran this on a small llms.txt: it extracts the title and the two documentation links. The format is simple Markdown, so it is easy to generate from your docs.
import re
llms = """# Acme Tasks
> Simple boards for small teams.
## Docs
- [Getting started](https://acme.example/docs/start): set up in 5 minutes
- [API](https://acme.example/docs/api): REST API reference
"""
links = re.findall(r"- \[(.+?)\]\((.+?)\)", llms)
print(links, llms.splitlines()[0])
Output:
[('Getting started', 'https://acme.example/docs/start'), ('API', 'https://acme.example/docs/api')] # Acme TasksPriorities: pages first
Good pages, working robots rules, a sitemap and consistent facts will help far more than any single file. Add llms.txt as a low-cost experiment after the basics are solid.
Quick check: What is the right way to think about llms.txt today?
- A legal requirement
- A guaranteed ranking factor
- A replacement for robots.txt
- A cheap experiment that does not replace good pages
Answer
A cheap experiment that does not replace good pages — Adoption is unproven, so keep expectations modest and do the fundamentals first.