Lesson 27 / 31

Evaluating Quality: Datasets, Metrics, Tracing

Measure answer and retrieval quality on realistic examples.

A golden set and traces

Build a dataset of 50 to 200 realistic questions with expected answers or source documents, including unanswerable ones. Measure retrieval (recall@k, MRR) and answers (correctness, faithfulness to the context, relevance, correct refusals, citation accuracy), using exact checks where possible and a calibrated LLM judge for open-ended text. Tools help: LangSmith datasets and evaluators, LlamaIndex's evaluation modules, and open-source options such as RAGAS. Tie every score to a trace so you can open the failing case and see the retrieved chunks, the final prompt and the model reply. Run the set in CI before releasing changes.

Quick check: Why link each evaluation score to a trace?

  • To avoid metrics
  • To hide failures
  • To remove datasets
  • To inspect retrieved chunks, the prompt and the reply for failing cases
Answer

To inspect retrieved chunks, the prompt and the reply for failing cases — A score says that something failed; a trace shows why.