Lesson 27 / 27

Revision: Cheat Sheet and Self-Check

Review the key ideas of the whole course.

Cheat sheet

Idea: retrieve relevant chunks, augment the prompt, generate a grounded, cited answer. Ingest: clean extraction, structure-aware chunks (200 to 500 tokens, small overlap), metadata for filters, permissions and citations. Retrieve: BM25 for exact terms, embeddings for meaning, hybrid with RRF, ANN indexes at scale. Improve: query rewriting, cross-encoder reranking, MMR for diversity, tuned k with a relevance threshold. Generate: numbered context, "answer only from context", verify citations in code, groundedness check, graceful refusal. Evaluate: golden set, recall@k and MRR, faithfulness, unanswerable questions, failure tally by stage. Production: incremental updates, access control, injection defence, latency and cost controls; add agentic or graph RAG only when measured need appears.

Quick check: Users search with "work from home" but the policy says "remote work". Which retrieval method helps most?

  • Lowering the temperature
  • Exact keyword match only
  • Increasing chunk overlap
  • Embedding (semantic) search, ideally combined with BM25
Answer

Embedding (semantic) search, ideally combined with BM25 — Embeddings bridge different wording; BM25 still covers exact terms.

Quick check: Which number best tells you whether retrieval is finding the evidence?

  • Average word length
  • Recall@k on a golden set
  • GPU temperature
  • Number of files
Answer

Recall@k on a golden set — Recall@k measures how often relevant chunks appear in the retrieved set.

Quick check: What is the safest way to enforce per-user document access?

  • Ask the model not to reveal secrets
  • Filter in the retrieval query so forbidden text is never fetched
  • Hide it in the UI
  • Rely on chunk size
Answer

Filter in the retrieval query so forbidden text is never fetched — If forbidden text reaches the prompt, it can leak.