Vector Databases

Learn to store, index, filter and operate vectors in real systems: pgvector with SQL, Qdrant and Chroma APIs, HNSW tuning, filtered and hybrid search, sharding, replication, capacity planning, security and choosing a database, with every example run for real.

Start course →

Syllabus

What a Vector Database Is

  1. Why Vectors Need a Database
  2. The Landscape: Libraries, Extensions, Dedicated Engines, Managed Services
  3. Data Model: Collections, Records, Payloads and Namespaces
  4. What to Expect: Approximate, Eventually Consistent, Not a General Database

pgvector: Vectors Inside PostgreSQL

  1. Creating a Table With a vector Column
  2. Distance Operators and ORDER BY
  3. Filtering With WHERE: SQL and Vectors Together
  4. Transactions and Types: vector, halfvec

Indexing in Practice

  1. Without an Index: The Exact Sequential Scan
  2. Creating an HNSW Index and Reading the Plan
  3. Measuring and Tuning Recall (ef_search)
  4. Index Size, Build Time and IVFFlat

Filtered and Hybrid Search

  1. The Filtered-Search Problem
  2. Pre-Filtering, Partial Indexes and Partitions
  3. Hybrid Search: Full-Text Plus Vectors
  4. Filtering in a Dedicated Engine: Qdrant

Other Engines and How to Compare Them

  1. Chroma: A Developer-Friendly Embedded Store
  2. Comparing Engines: A Practical Checklist
  3. Benchmarking Honestly

Operating at Scale

  1. Bulk Ingestion and Batching
  2. Sharding: Splitting Data Across Machines
  3. Replication, Quorums and Read-Your-Writes
  4. Capacity Planning: Memory Is the Budget

Security, Backups and Choosing

  1. Security and Multi-Tenancy
  2. Backups, Migrations and Re-Embedding
  3. Monitoring and Cost

Putting It Together

  1. Case Study: Choosing and Running a Store for a SaaS Knowledge Base
  2. Revision: Cheat Sheet and Self-Check