Lesson 18 / 28

Error Analysis and Improving Quality

Find why queries fail and pick the right fix.

Diagnose before you tune

For each failing query ask in order: Is the answer in the corpus at all? Was it chunked sensibly (not cut mid-idea, not buried in a huge chunk)? Does the text use different vocabulary (add keyword/hybrid search, query rewriting, or a better or domain-adapted model)? Is the right chunk retrieved but ranked too low (add a reranker)? Is the embedding model weak for the language or domain (try another, or fine-tune on your own question-passage pairs with hard negatives)? Is the ANN index losing recall (compare with flat; raise nprobe/efSearch)? Are filters removing the answer? Tally failures by cause across the whole golden set and fix the biggest bucket first.

Failure tally and first fixes

An example of how a team might sort 40 failed queries. The numbers are illustrative, not measured.

Cause                                     Count   First fix to try
answer not in the corpus                    3     add content / return "not found"
bad chunking (idea split / buried)          7     structure-aware chunks, titles in chunks
vocabulary mismatch (exact words differ)   11     hybrid search, query rewrite, better model
right chunk found but ranked low           10     cross-encoder reranker
ANN index missed it (flat finds it)         4     raise nprobe / efSearch, check quantisation
filter removed the answer                   5     check metadata, over-fetch before filtering

Quick check: A relevant chunk is retrieved but ranked 14th. What is the most direct fix?

  • Add a reranker to reorder the shortlist
  • Delete the filter
  • Double the dimensions
  • Ignore it
Answer

Add a reranker to reorder the shortlist — The evidence is found, so improve ordering rather than recall.