Lesson 20 / 28

Where Vectors Live

Compare libraries, extensions to databases you already run, and dedicated vector databases.

Library, extension or database

Three broad choices. (1) An in-process library such as FAISS: very fast, great for experiments and read-mostly data, but you handle persistence, replication, filtering and updates yourself. (2) A vector extension for a database you already run (for example pgvector for PostgreSQL, or vector search in Elasticsearch/OpenSearch, MongoDB or Redis): vectors sit next to your relational data and metadata, you reuse backups, access control and transactions, and filtering with SQL is natural; performance is good to large scale but may trail specialised engines. (3) A dedicated vector database (such as Qdrant, Milvus, Weaviate, Pinecone): built for scale, filtering, hybrid search, replication and operations, at the cost of another system to run or pay for. Start with what you already operate; move to a specialised engine when measured needs (scale, latency, features) demand it. The next course in this series compares vector databases in detail.

Store, update, scale, protect

Choose where vectors live, keep them fresh, control cost and protect the data they came from.

Four concerns: storage, freshness, scale, security.
Figure 6.1 — Storage, freshness, scale and security.

Where to start

A rough guide; measure before you commit.

Situation                                        Reasonable start
experiment / notebook / offline batch              FAISS or numpy (flat)
already on PostgreSQL, < ~ tens of millions        pgvector (+ SQL filters)
already on Elasticsearch/OpenSearch, want hybrid   its vector + BM25 search
large scale, heavy filtering, many tenants         dedicated vector database
unsure                                             start simple; keep an abstraction so you can move

Hide the store behind an interface

A small wrapper with add, delete and search lets you change engines later without rewriting the app.

Quick check: What is an advantage of a vector extension in a database you already run?

  • It is always faster than every other option
  • Vectors sit with your data, reusing backups, access control and SQL filters
  • It needs no storage
  • It removes the need for embeddings
Answer

Vectors sit with your data, reusing backups, access control and SQL filters — One system to operate often beats two, until scale demands otherwise.