Lesson 20 / 31

Building a Vector Index and Retriever

Embed nodes and retrieve the most similar ones.

VectorStoreIndex in a few lines

VectorStoreIndex.from_documents(docs) parses the Documents into nodes, embeds them with the configured embedding model and stores them (in memory by default, or in an external vector store). Global defaults for the LLM and embedding model live in Settings; set them explicitly so behaviour is clear. index.as_retriever(similarity_top_k=2) returns a retriever whose retrieve(query) gives NodeWithScore objects: the node plus a similarity score. Other index types exist (summary, keyword table, knowledge graph, property graph) for different query patterns.

Retrieval with a toy embedding, run

I ran this offline in a Python virtual environment with langchain-core 1.6.6, langchain-text-splitters 1.1.2 and llama-index-core 0.14.25. No API key or network call is needed because a fake model or a toy embedding stands in for the real one. A custom BaseEmbedding hashes words into 16 buckets (no meaning, so it only matches shared words). The leave question scores 0.647 against the leave document and 0.316 against the hotel document, which shares the word "per". A real embedding model would also match synonyms.

from llama_index.core import Document, VectorStoreIndex, Settings
from llama_index.core.embeddings import BaseEmbedding
from llama_index.core.llms import MockLLM
import hashlib, math

class HashEmbedding(BaseEmbedding):
    """Toy deterministic embedding: hashes words into 16 buckets (no meaning, like a bag of words)."""
    def _vec(self, text):
        v = [0.0] * 16
        for w in text.lower().split():
            v[int(hashlib.md5(w.strip(".,?").encode()).hexdigest(), 16) % 16] += 1
        n = math.sqrt(sum(x * x for x in v)) or 1
        return [x / n for x in v]
    def _get_query_embedding(self, query): return self._vec(query)
    def _get_text_embedding(self, text): return self._vec(text)
    async def _aget_query_embedding(self, query): return self._vec(query)

Settings.llm = MockLLM()
Settings.embed_model = HashEmbedding()

docs = [Document(text="Employees get 24 days of paid leave per year."),
        Document(text="Hotels are capped at 6000 rupees per night."),
        Document(text="Report lost laptops within 24 hours.")]
index = VectorStoreIndex.from_documents(docs)
retriever = index.as_retriever(similarity_top_k=2)
for r in retriever.retrieve("how many days of leave per year"):
    print(round(r.score, 3), "|", r.node.get_content())

Output:

0.647 | Employees get 24 days of paid leave per year.
0.316 | Hotels are capped at 6000 rupees per night.

Set Settings explicitly

If you rely on defaults, the library may try to call a hosted model and ask for an API key. Setting Settings.llm and Settings.embed_model makes your app predictable.

Quick check: What does a retriever's result include besides the node text?

  • A GPU temperature
  • A similarity score
  • A password
  • A new model
Answer

A similarity score — NodeWithScore pairs each node with how well it matched.