Lesson 2 / 25

Common Failure Modes

Recognise hallucination, harmful output, privacy leaks, bias, misuse and over-reliance.

Six things that go wrong

Hallucination: a fluent but false or invented answer. Harmful content: abusive, dangerous or inappropriate text. Privacy leaks: personal data repeated from a prompt, a document or another user's session. Bias: worse results or tone for some groups. Misuse: someone steering the system to do something it should not (injection, jailbreaks, spam). Over-reliance: users trusting answers without checking. Each needs a different safeguard, so naming the failure mode is the first step.

Failure mode to safeguard

A starting map. Real systems use several safeguards per row.

Failure mode        First safeguards
Hallucination       ground answers in documents, cite sources, say "I don't know"
Harmful content     moderation on input and output, clear policies
Privacy leak        minimise data, redact PII, isolate user sessions
Bias                test across groups, review failures, human review
Misuse              least-privilege tools, rate limits, injection defences
Over-reliance       show uncertainty, link evidence, human sign-off on high stakes

Quick check: A model confidently states a policy that does not exist. Which failure mode is this?

  • Hallucination
  • Rate limiting
  • Caching
  • Compression
Answer

Hallucination — A fluent invented answer is the classic definition of hallucination.