Lesson 18 / 25

A RAG Pipeline in n8n

Build ingestion and query workflows with loaders, splitters, embeddings and a vector store.

Two workflows

RAG (retrieval-augmented generation) answers from your documents. It needs two flows. Ingestion (run when documents change): load files, split them into chunks, turn chunks into embeddings, and store them in a vector store (n8n supports options such as Qdrant, Pinecone, Supabase and an in-memory store for testing). Query: the agent gets the vector store as a tool, retrieves the best chunks for the question and answers using them.

The two flows side by side

The same embedding model must be used on both sides, or the vectors will not match.

INGEST   Trigger -> Load file -> Split text (500-1000 chars, ~10% overlap)
                   -> Embeddings (model X) -> Vector store: insert

QUERY    Chat trigger -> AI Agent
                          |- Chat model
                          |- Memory
                          \- Tool: Vector store (retrieve, Embeddings: model X)

Tell the agent to cite and admit gaps

In the system message, require answers to quote or name the source chunk and to say "I could not find that in the documents" when retrieval returns nothing relevant. Retrieved text is data, not instructions.

Quick check: Why must ingestion and query use the same embedding model?

  • Vectors from different models are not comparable
  • It saves disk space only
  • n8n forbids two models
  • It is only a naming rule
Answer

Vectors from different models are not comparable — Similarity search only works when both sides use the same vector space.