Lesson 22 / 31

Persisting and Reloading an Index

Avoid re-embedding on every start.

Embed once, load many times

Embedding a large corpus costs time and money, so save the index. index.storage_context.persist(persist_dir=...) writes the document store, index store and vector store as files, and load_index_from_storage(StorageContext.from_defaults(persist_dir=...)) restores them. For production, use a real vector database (Qdrant, Pinecone, pgvector, Chroma and others) through the corresponding integration so data survives restarts and scales. Update indexes incrementally when documents change (insert, delete or refresh by document ID) rather than rebuilding everything, and remember that changing the embedding model requires re-embedding.

Persist and reload, run

I ran this offline in a Python virtual environment with langchain-core 1.6.6, langchain-text-splitters 1.1.2 and llama-index-core 0.14.25. No API key or network call is needed because a fake model or a toy embedding stands in for the real one. Persisting writes JSON files for the stores (the exact list can differ between versions), and loading them back restores the index with its 1 document.

import tempfile, os
from llama_index.core import Document, VectorStoreIndex, StorageContext, load_index_from_storage, Settings
from llama_index.core.embeddings import MockEmbedding
from llama_index.core.llms import MockLLM

Settings.llm = MockLLM(); Settings.embed_model = MockEmbedding(embed_dim=8)
index = VectorStoreIndex.from_documents([Document(text="Persisted index demo.")])
with tempfile.TemporaryDirectory() as d:
    index.storage_context.persist(persist_dir=d)
    print(sorted(os.listdir(d)))
    again = load_index_from_storage(StorageContext.from_defaults(persist_dir=d))
    print("reloaded docs:", len(again.docstore.docs))

Output:

['default__vector_store.json', 'docstore.json', 'graph_store.json', 'image__vector_store.json', 'index_store.json']
reloaded docs: 1

Quick check: Why persist an index?

  • To remove the retriever
  • To make queries wrong
  • To hide metadata
  • To avoid paying to re-embed the corpus at every start
Answer

To avoid paying to re-embed the corpus at every start — Embedding is the expensive step, so reuse its result.