Lesson 22 / 27

Keeping the Index Fresh

Update, delete and version documents without rebuilding everything.

Stale answers are wrong answers

Documents change, so the index must too. Use incremental updates: detect changed files by content hash or modified time, re-chunk and re-embed only those, and delete the old chunks (track them by document ID). Handle deletions and access changes promptly, since a removed or restricted document must stop being retrievable. Store version and effective dates as metadata and prefer the latest valid version in filters. When you change the embedding model or chunking, build a new index alongside the old, compare on the golden set, then switch over.

Fresh, safe, fast, smart

Keep the index current, protect data, control cost and consider more advanced patterns.

Four concerns: freshness, security, performance, advanced patterns.
Figure 7.1 — Freshness, security, performance and advanced patterns.

Schedule a full rebuild too

Incremental updates can drift. A periodic full rebuild into a new index, verified against the golden set, resets accumulated errors.

Quick check: How should you roll out a new embedding model?

  • Change nothing
  • Mix old and new vectors in one index
  • Delete the old index first
  • Build a new index alongside the old, compare, then switch
Answer

Build a new index alongside the old, compare, then switch — Vectors from different models are incomparable, and a side-by-side check avoids regressions.