Lesson 18 / 31

Making Retrieval Better: Filters, MMR, Reranking

Improve what reaches the prompt.

Quality in, quality out

Most RAG failures are retrieval failures. Improvements available in LangChain and its integrations: metadata filters to restrict by tenant, date or permissions; MMR (search_type="mmr") to diversify near-duplicate results; hybrid search combining keyword and vector scores where the store supports it; multi-query retrieval that asks the model to rewrite the question several ways; contextual compression and rerankers that re-order or trim retrieved passages; and a parent-document retriever that matches small chunks but returns larger sections. Measure with a set of real questions (recall@k) before and after each change.

Change one thing, re-measure

Chunk size, k, embeddings and reranking interact. Test them one at a time on the same questions.

Quick check: What does MMR add to retrieval?

  • More diverse results by reducing near-duplicates
  • Faster GPUs
  • Free embeddings
  • Automatic fine-tuning
Answer

More diverse results by reducing near-duplicates — Maximal marginal relevance balances relevance with diversity.