# Evaluating Answers: Faithfulness and Relevance — Retrieval-Augmented Generation (RAG)

Source: https://www.geekswithgeeks.com/en/rag/e-answers

> Score the generated answers on correctness, support and usefulness.

## Beyond "looks good"

Score each answer on several axes: **correctness** against a reference answer, **faithfulness/groundedness** to the retrieved context, **answer relevance** to the question, **completeness**, correct **refusals** on questions the documents cannot answer (include some in your set), and **citation accuracy**. Rubric scoring by an **LLM judge** scales well, but calibrate it against human ratings on a sample, watch for judge bias (favouring long answers) and keep the judge prompts fixed. Frameworks such as RAGAS or TruLens offer ready metrics; use them as aids, not as ground truth. Add latency, cost and user feedback (thumbs up/down) to the dashboard.

## Include unanswerable questions

Add 10 to 20% questions whose answer is not in the documents. The correct behaviour is a refusal, and many systems fail exactly here.

**Quiz:** Why include unanswerable questions in the test set?

- [ ] To make the set shorter
- [x] To check that the system refuses instead of inventing answers
- [ ] To lower recall
- [ ] To avoid judges

*Answer:* To check that the system refuses instead of inventing answers. Refusal behaviour is only visible if the set contains questions that should be refused.
