Lesson 27 / 27
Revision: Cheat Sheet and Self-Check
Review the key ideas of the whole course.
Cheat sheet
Idea: retrieve relevant chunks, augment the prompt, generate a grounded, cited answer. Ingest: clean extraction, structure-aware chunks (200 to 500 tokens, small overlap), metadata for filters, permissions and citations. Retrieve: BM25 for exact terms, embeddings for meaning, hybrid with RRF, ANN indexes at scale. Improve: query rewriting, cross-encoder reranking, MMR for diversity, tuned k with a relevance threshold. Generate: numbered context, "answer only from context", verify citations in code, groundedness check, graceful refusal. Evaluate: golden set, recall@k and MRR, faithfulness, unanswerable questions, failure tally by stage. Production: incremental updates, access control, injection defence, latency and cost controls; add agentic or graph RAG only when measured need appears.
Quick check: Users search with "work from home" but the policy says "remote work". Which retrieval method helps most?
- Lowering the temperature
- Exact keyword match only
- Increasing chunk overlap
- Embedding (semantic) search, ideally combined with BM25
Answer
Embedding (semantic) search, ideally combined with BM25 — Embeddings bridge different wording; BM25 still covers exact terms.
Quick check: Which number best tells you whether retrieval is finding the evidence?
- Average word length
- Recall@k on a golden set
- GPU temperature
- Number of files
Answer
Recall@k on a golden set — Recall@k measures how often relevant chunks appear in the retrieved set.
Quick check: What is the safest way to enforce per-user document access?
- Ask the model not to reveal secrets
- Filter in the retrieval query so forbidden text is never fetched
- Hide it in the UI
- Rely on chunk size
Answer
Filter in the retrieval query so forbidden text is never fetched — If forbidden text reaches the prompt, it can leak.