# Bulk Ingestion and Batching — Vector Databases

Source: https://www.geekswithgeeks.com/en/vector-databases/o-ingest

> Load millions of vectors efficiently and idempotently.

## Batches, retries and ordering of work

Inserting vectors one request at a time wastes most of the time on network round trips. **Batch** writes (hundreds to a few thousand records per request, within the engine's limits), use **parallel workers** with bounded concurrency, and make writes **idempotent** (upsert by deterministic ID) so retries after a failure cannot create duplicates. Generating **embeddings** is usually the slowest and costliest step, so cache by content hash and embed in batches too. For a **first big load**, create the table or collection, load data, and **build the ANN index afterwards** where the engine allows it (faster than inserting into an existing index), then verify counts and sample recall. For **ongoing** updates, use a queue or change feed so ingestion can be retried and monitored, and track **lag** (time from change to searchable).

## Ingest, shard, replicate, plan

Batch your writes, spread data across shards, replicate for availability, and do the capacity arithmetic before you buy.

![Four jobs: ingest, shard, replicate, plan.](assets/figures/vector-databases/section-6-map.svg) — Figure 6.1 — Ingest, shard, replicate and plan.

## Why batching matters, run

I ran this plain-Python (standard library only) example. This is a model with assumed numbers (5 ms per request and 0.2 ms of work per vector), not a measurement of any database. Loading 100,000 vectors one by one is dominated by round trips (520 s), while batches of 1,000 cut it to about 20 s.

```python
import time
batch_sizes = [1, 10, 100, 1000]
calls_overhead = 0.005                      # assume 5 ms of network round-trip per request (illustrative)
per_item = 0.0002                           # assume 0.2 ms of work per vector (illustrative)
n = 100_000
for b in batch_sizes:
    requests = -(-n // b)
    total = requests * calls_overhead + n * per_item
    print(f"batch={b:5d} requests={requests:7d} estimated time={total:8.1f} s")

```

Output:

```
batch=    1 requests= 100000 estimated time=   520.0 s
batch=   10 requests=  10000 estimated time=    70.0 s
batch=  100 requests=   1000 estimated time=    25.0 s
batch= 1000 requests=    100 estimated time=    20.5 s
```

## Embed once, cache by content hash

Re-embedding unchanged text wastes money. Skip chunks whose hash did not change.

**Quiz:** Why make writes idempotent with deterministic IDs?

- [ ] To make vectors bigger
- [x] So retries after failures cannot create duplicates
- [ ] To avoid indexes
- [ ] To speed up embeddings

*Answer:* So retries after failures cannot create duplicates. Upserting the same ID twice leaves one record.
