# Evaluating Quality: Datasets, Metrics, Tracing — LangChain / LlamaIndex

Source: https://www.geekswithgeeks.com/en/langchain-llamaindex/p-eval

> Measure answer and retrieval quality on realistic examples.

## A golden set and traces

Build a **dataset** of 50 to 200 realistic questions with expected answers or source documents, including unanswerable ones. Measure **retrieval** (recall@k, MRR) and **answers** (correctness, faithfulness to the context, relevance, correct refusals, citation accuracy), using exact checks where possible and a calibrated **LLM judge** for open-ended text. Tools help: **LangSmith** datasets and evaluators, LlamaIndex's evaluation modules, and open-source options such as RAGAS. Tie every score to a **trace** so you can open the failing case and see the retrieved chunks, the final prompt and the model reply. Run the set in CI before releasing changes.

**Quiz:** Why link each evaluation score to a trace?

- [ ] To avoid metrics
- [ ] To hide failures
- [ ] To remove datasets
- [x] To inspect retrieved chunks, the prompt and the reply for failing cases

*Answer:* To inspect retrieved chunks, the prompt and the reply for failing cases. A score says that something failed; a trace shows why.
