Lesson 26 / 28
Embeddings for RAG and Agent Memory
Ground language models in your data with retrieval.
Retrieval is the first half of RAG
In retrieval-augmented generation (RAG) the embedding index finds the passages relevant to a question, and the language model answers from them, with citations. The answer can be no better than the evidence retrieved, so retrieval quality (recall@k) is the main lever; everything in this course applies: good chunks and metadata, a suitable embedding model, hybrid search, reranking, filters and evaluation. Agent memory uses the same idea: store past facts, conversations or tool results as vectors and retrieve the few relevant to the current task instead of stuffing the whole history into the prompt. Remember to let the system say "not found" when the best match is weak, to cite sources, and to treat retrieved text as untrusted.
Quick check: What is the main limit on a RAG answer's quality?
- The relevance of the evidence that retrieval finds
- The colour of the UI
- The number of vector dimensions alone
- The length of the API key
Answer
The relevance of the evidence that retrieval finds — If the evidence is missing or wrong, the model cannot answer faithfully.