# Data Model: Collections, Records, Payloads and Namespaces — Vector Databases

Source: https://www.geekswithgeeks.com/en/vector-databases/b-datamodel

> Design what to store with each vector.

## Decide the record before the database

Most engines group records into **collections** (or tables/indexes) where every vector has the **same dimension and distance metric**, tied to one embedding model. Design the **record**: a stable **ID** (derived from document ID and chunk number, so re-ingesting overwrites rather than duplicates); the **vector**; the **text** (or a pointer to it); and **payload fields** you will filter or display on (tenant, language, year, source, page, permissions, content hash, embedding model name). Index the payload fields you filter on. Decide how to separate customers: a **collection per tenant** (strong isolation, many small indexes), **one collection with a tenant field and filter** (efficient, relies on filters being enforced), or **namespaces/partitions** where offered. Keep large documents out of the vector store when you can: store the text elsewhere and keep an ID.

## Deterministic IDs prevent duplicates, run

I ran this plain-Python (standard library only) example. The same document and chunk number always produce the same ID, so re-ingesting after an edit overwrites the old record instead of adding a duplicate. A changed text gives a different content hash, which tells you the vector needs recomputing.

```python
import hashlib

def record_id(doc_id, chunk_no): return f"{doc_id}#chunk-{chunk_no}"
def content_hash(text): return hashlib.sha256(text.encode()).hexdigest()[:12]

v1 = "Unused leave up to 5 days can be carried over."
v2 = "Unused leave up to 10 days can be carried over."
print(record_id("hr-policy-2025", 3), record_id("hr-policy-2025", 3))
print("hash v1:", content_hash(v1), "| hash v2:", content_hash(v2), "| changed:", content_hash(v1) != content_hash(v2))
```

Output:

```
hr-policy-2025#chunk-3 hr-policy-2025#chunk-3
hash v1: 1c41c4893eab | hash v2: 20718f62c63d | changed: True
```

## Store the model name with every record

A `model` field makes mixed-model mistakes detectable and re-embedding auditable.

**Quiz:** Why derive record IDs from document ID and chunk number?

- [ ] Databases forbid other IDs
- [ ] IDs must be random
- [ ] It makes vectors smaller
- [x] Re-ingesting overwrites instead of duplicating

*Answer:* Re-ingesting overwrites instead of duplicating. Stable IDs make updates idempotent.
