Lesson 16 / 31

Embeddings, Vector Stores and Retrievers

Index chunks and search them by similarity.

Store vectors, search by closeness

An embeddings model turns text into vectors. A vector store saves the vectors with their Documents and supports similarity_search. A retriever is a runnable that takes a query and returns Documents; store.as_retriever(search_kwargs={"k": 3}) wraps a store. LangChain has a common interface over many stores (in-memory, FAISS, Chroma, pgvector, Pinecone, Qdrant and others) so you can start in memory and move to a database later. Use metadata filters for hard constraints (tenant, language, permissions). The demo below uses a deterministic fake embedding, which has no meaning: it only shows the API, and a text is most similar to itself.

An in-memory vector store, run

I ran this offline in a Python virtual environment with langchain-core 1.6.6, langchain-text-splitters 1.1.2 and llama-index-core 0.14.25. No API key or network call is needed because a fake model or a toy embedding stands in for the real one. Searching with the exact text of the hotel document ranks it first (the travel source). Because the fake embedding carries no meaning, the second hit is not semantically related; with a real embedding model the ranking would reflect meaning. The retriever returns the single best match for the laptop text.

from langchain_core.documents import Document
from langchain_core.vectorstores import InMemoryVectorStore
from langchain_core.embeddings import DeterministicFakeEmbedding

docs = [
    Document(page_content="Employees get 24 days of paid leave.", metadata={"src": "hr"}),
    Document(page_content="Hotels are capped at 6000 rupees per night.", metadata={"src": "travel"}),
    Document(page_content="Report lost laptops within 24 hours.", metadata={"src": "security"}),
]
store = InMemoryVectorStore(DeterministicFakeEmbedding(size=32))
store.add_documents(docs)

# The fake embedding is deterministic but has no meaning, so the same text always matches itself best.
hits = store.similarity_search("Hotels are capped at 6000 rupees per night.", k=2)
print([h.metadata["src"] for h in hits])
retriever = store.as_retriever(search_kwargs={"k": 1})
print(retriever.invoke("Report lost laptops within 24 hours.")[0].page_content)

Output:

['travel', 'hr']
Report lost laptops within 24 hours.

Quick check: What is a retriever?

  • A prompt template
  • A type of GPU
  • A runnable that takes a query and returns relevant Documents
  • A model weight file
Answer

A runnable that takes a query and returns relevant Documents — Retrievers hide the search backend behind a simple interface.