Lesson 20 / 28

Bulk Ingestion and Batching

Load millions of vectors efficiently and idempotently.

Batches, retries and ordering of work

Inserting vectors one request at a time wastes most of the time on network round trips. Batch writes (hundreds to a few thousand records per request, within the engine's limits), use parallel workers with bounded concurrency, and make writes idempotent (upsert by deterministic ID) so retries after a failure cannot create duplicates. Generating embeddings is usually the slowest and costliest step, so cache by content hash and embed in batches too. For a first big load, create the table or collection, load data, and build the ANN index afterwards where the engine allows it (faster than inserting into an existing index), then verify counts and sample recall. For ongoing updates, use a queue or change feed so ingestion can be retried and monitored, and track lag (time from change to searchable).

Ingest, shard, replicate, plan

Batch your writes, spread data across shards, replicate for availability, and do the capacity arithmetic before you buy.

Four jobs: ingest, shard, replicate, plan.
Figure 6.1 — Ingest, shard, replicate and plan.

Why batching matters, run

I ran this plain-Python (standard library only) example. This is a model with assumed numbers (5 ms per request and 0.2 ms of work per vector), not a measurement of any database. Loading 100,000 vectors one by one is dominated by round trips (520 s), while batches of 1,000 cut it to about 20 s.

import time
batch_sizes = [1, 10, 100, 1000]
calls_overhead = 0.005                      # assume 5 ms of network round-trip per request (illustrative)
per_item = 0.0002                           # assume 0.2 ms of work per vector (illustrative)
n = 100_000
for b in batch_sizes:
    requests = -(-n // b)
    total = requests * calls_overhead + n * per_item
    print(f"batch={b:5d} requests={requests:7d} estimated time={total:8.1f} s")

Output:

batch=    1 requests= 100000 estimated time=   520.0 s
batch=   10 requests=  10000 estimated time=    70.0 s
batch=  100 requests=   1000 estimated time=    25.0 s
batch= 1000 requests=    100 estimated time=    20.5 s

Embed once, cache by content hash

Re-embedding unchanged text wastes money. Skip chunks whose hash did not change.

Quick check: Why make writes idempotent with deterministic IDs?

  • To make vectors bigger
  • So retries after failures cannot create duplicates
  • To avoid indexes
  • To speed up embeddings
Answer

So retries after failures cannot create duplicates — Upserting the same ID twice leaves one record.