Lesson 4 / 27

What LLMs Are Good and Bad At

Know the strengths and typical failure modes.

Fluent is not the same as correct

LLMs are strong at language tasks: summarising, rewriting, translating, classifying, extracting fields, explaining, drafting code and following examples. Typical weaknesses: hallucination (confident but invented facts or citations), knowledge cutoff (nothing after training unless supplied), unreliable exact arithmetic and counting, inconsistency between runs, sensitivity to wording, and limited memory (only what fits in the context window). The practical rule: use LLMs where a wrong answer is cheap to detect or can be verified, and add retrieval, tools and checks where it is not.

Verify what you cannot afford to get wrong

Ask yourself: if this answer is wrong, who is hurt and how fast would I notice? If the answer is "badly and slowly", add a check or a human.

Quick check: What is hallucination?

  • A GPU overheating
  • Confident output that is false or invented
  • A slow network
  • Compressing a prompt
Answer

Confident output that is false or invented — The model generates plausible text, not verified facts.