Lesson 3 / 27
RAG vs Fine-Tuning vs Long Context
Choose the right approach for a knowledge problem.
Three tools, three jobs
RAG suits large or changing knowledge, per-user permissions and answers that must cite sources. Fine-tuning changes the model's behaviour (style, format, narrow skills) and is a poor way to store facts that change. Long context (paste everything into the prompt) is simplest for small, stable document sets, but costs more per call, can slow responses, and models may miss details buried in the middle of very long inputs. These combine: a fine-tuned model can be used with RAG, and a long window lets you retrieve more chunks. Start with the simplest approach that meets your accuracy, cost and freshness needs.
A decision table
Starting points; confirm with your own tests.
Situation Lean toward
20 pages, rarely change, low traffic long context
10,000 docs, change weekly, need citations RAG
Need a consistent tone/format on every reply fine-tuning (+ RAG for facts)
Different users may see different documents RAG with permission filters
Unsure prototype RAG, measure on 50 real questionsCombine approaches
A small, stable "style guide" can live in the system prompt while changing facts come from retrieval. Use each tool for what it does best.
Quick check: Which need points most clearly to RAG?
- Making the model smaller
- A fixed brand tone only
- A tiny static FAQ of 5 lines
- Answers from thousands of changing documents, with citations
Answer
Answers from thousands of changing documents, with citations — Large, changing, citable knowledge is RAG's sweet spot.