Lesson 20 / 27
Evaluating Answers: Faithfulness and Relevance
Score the generated answers on correctness, support and usefulness.
Beyond "looks good"
Score each answer on several axes: correctness against a reference answer, faithfulness/groundedness to the retrieved context, answer relevance to the question, completeness, correct refusals on questions the documents cannot answer (include some in your set), and citation accuracy. Rubric scoring by an LLM judge scales well, but calibrate it against human ratings on a sample, watch for judge bias (favouring long answers) and keep the judge prompts fixed. Frameworks such as RAGAS or TruLens offer ready metrics; use them as aids, not as ground truth. Add latency, cost and user feedback (thumbs up/down) to the dashboard.
Include unanswerable questions
Add 10 to 20% questions whose answer is not in the documents. The correct behaviour is a refusal, and many systems fail exactly here.
Quick check: Why include unanswerable questions in the test set?
- To make the set shorter
- To check that the system refuses instead of inventing answers
- To lower recall
- To avoid judges
Answer
To check that the system refuses instead of inventing answers — Refusal behaviour is only visible if the set contains questions that should be refused.