Lesson 18 / 27
Groundedness Checks and Refusals
Detect unsupported answers and refuse gracefully.
Supported by the evidence?
Groundedness (also called faithfulness) asks whether every claim in the answer is supported by the retrieved text. Cheap checks compare the answer's content words with the context: low overlap is a warning sign. Stronger checks use an LLM judge or an NLI model to test each claim against the context. When groundedness is low or retrieval confidence is below the threshold, refuse gracefully ("I couldn't find this in the documents"), offer related sources or hand off to a human. A correct refusal is better than a confident guess.
A word-overlap groundedness heuristic, run
I ran this plain-Python (standard library only) example. The faithful answer shares 80% of its content words with the context; the answer inventing "20 days and cash them out" shares only 25% and is flagged. This is a crude heuristic, not a substitute for a proper judge.
import re
def words(t): return set(re.findall(r"[a-z0-9]+", t.lower())) - {"the","a","an","of","to","is","are","can","be","you","per","up"}
context = "Unused leave up to 5 days can be carried over to the next year."
for ans in ("You can carry over 5 days of leave.", "You can carry over 20 days and cash them out."):
aw = words(ans); support = len(aw & words(context)) / len(aw)
print(round(support, 2), "grounded" if support >= 0.6 else "CHECK", "|", ans)
Output:
0.8 grounded | You can carry over 5 days of leave. 0.25 CHECK | You can carry over 20 days and cash them out.
Show confidence honestly
When the evidence is partial, say which part is supported and which is not, instead of presenting a complete-sounding answer.
Quick check: What should the system do when evidence is too weak?
- Say it could not find the answer in the documents
- Invent a plausible answer
- Increase the temperature
- Delete the index
Answer
Say it could not find the answer in the documents — An honest "not found" protects users from confident errors.