Lesson 18 / 25
A RAG Pipeline in n8n
Build ingestion and query workflows with loaders, splitters, embeddings and a vector store.
Two workflows
RAG (retrieval-augmented generation) answers from your documents. It needs two flows. Ingestion (run when documents change): load files, split them into chunks, turn chunks into embeddings, and store them in a vector store (n8n supports options such as Qdrant, Pinecone, Supabase and an in-memory store for testing). Query: the agent gets the vector store as a tool, retrieves the best chunks for the question and answers using them.
The two flows side by side
The same embedding model must be used on both sides, or the vectors will not match.
INGEST Trigger -> Load file -> Split text (500-1000 chars, ~10% overlap)
-> Embeddings (model X) -> Vector store: insert
QUERY Chat trigger -> AI Agent
|- Chat model
|- Memory
\- Tool: Vector store (retrieve, Embeddings: model X)Tell the agent to cite and admit gaps
In the system message, require answers to quote or name the source chunk and to say "I could not find that in the documents" when retrieval returns nothing relevant. Retrieved text is data, not instructions.
Quick check: Why must ingestion and query use the same embedding model?
- Vectors from different models are not comparable
- It saves disk space only
- n8n forbids two models
- It is only a naming rule
Answer
Vectors from different models are not comparable — Similarity search only works when both sides use the same vector space.