Lesson 2 / 25
Common Failure Modes
Recognise hallucination, harmful output, privacy leaks, bias, misuse and over-reliance.
Six things that go wrong
Hallucination: a fluent but false or invented answer. Harmful content: abusive, dangerous or inappropriate text. Privacy leaks: personal data repeated from a prompt, a document or another user's session. Bias: worse results or tone for some groups. Misuse: someone steering the system to do something it should not (injection, jailbreaks, spam). Over-reliance: users trusting answers without checking. Each needs a different safeguard, so naming the failure mode is the first step.
Failure mode to safeguard
A starting map. Real systems use several safeguards per row.
Failure mode First safeguards
Hallucination ground answers in documents, cite sources, say "I don't know"
Harmful content moderation on input and output, clear policies
Privacy leak minimise data, redact PII, isolate user sessions
Bias test across groups, review failures, human review
Misuse least-privilege tools, rate limits, injection defences
Over-reliance show uncertainty, link evidence, human sign-off on high stakesQuick check: A model confidently states a policy that does not exist. Which failure mode is this?
- Hallucination
- Rate limiting
- Caching
- Compression
Answer
Hallucination — A fluent invented answer is the classic definition of hallucination.