Lesson 21 / 27
Failure Analysis: Where Did It Break?
Diagnose a bad answer by walking back through the pipeline.
Walk the pipeline backwards
For each wrong answer ask in order: (1) Is the answer in the documents at all? (2) Was it extracted and chunked intact? (3) Was the right chunk retrieved (in the top k)? (4) Was it ranked high enough to be sent? (5) Did the prompt present it clearly? (6) Did the model use it faithfully? Tally failures by stage over your golden set; the biggest bucket tells you where to invest. Typical fixes: extraction bugs, chunk size, hybrid search, query rewriting, reranking, better instructions, a stronger model.
A failure tally
An example of how a team might tally 40 failed questions to decide what to fix first. The numbers are illustrative, not measured.
Failure stage Count Fix to try first
answer not in the documents 4 add missing docs / better refusal
bad extraction (tables, scans) 6 improve parsing / OCR
right chunk not in top-k 18 hybrid search, chunking, rewrite
retrieved but ranked too low 7 reranker
retrieved and ranked, model misread 5 prompt, citations, stronger modelQuick check: What is the first question when an answer is wrong?
- Is the model too small?
- Is the answer actually in the documents?
- Is the font readable?
- Is the server red?
Answer
Is the answer actually in the documents? — If the knowledge is missing, no retrieval or prompt tweak can fix it.