Lesson 3 / 28

Data Model: Collections, Records, Payloads and Namespaces

Design what to store with each vector.

Decide the record before the database

Most engines group records into collections (or tables/indexes) where every vector has the same dimension and distance metric, tied to one embedding model. Design the record: a stable ID (derived from document ID and chunk number, so re-ingesting overwrites rather than duplicates); the vector; the text (or a pointer to it); and payload fields you will filter or display on (tenant, language, year, source, page, permissions, content hash, embedding model name). Index the payload fields you filter on. Decide how to separate customers: a collection per tenant (strong isolation, many small indexes), one collection with a tenant field and filter (efficient, relies on filters being enforced), or namespaces/partitions where offered. Keep large documents out of the vector store when you can: store the text elsewhere and keep an ID.

Deterministic IDs prevent duplicates, run

I ran this plain-Python (standard library only) example. The same document and chunk number always produce the same ID, so re-ingesting after an edit overwrites the old record instead of adding a duplicate. A changed text gives a different content hash, which tells you the vector needs recomputing.

import hashlib

def record_id(doc_id, chunk_no): return f"{doc_id}#chunk-{chunk_no}"
def content_hash(text): return hashlib.sha256(text.encode()).hexdigest()[:12]

v1 = "Unused leave up to 5 days can be carried over."
v2 = "Unused leave up to 10 days can be carried over."
print(record_id("hr-policy-2025", 3), record_id("hr-policy-2025", 3))
print("hash v1:", content_hash(v1), "| hash v2:", content_hash(v2), "| changed:", content_hash(v1) != content_hash(v2))

Output:

hr-policy-2025#chunk-3 hr-policy-2025#chunk-3
hash v1: 1c41c4893eab | hash v2: 20718f62c63d | changed: True

Store the model name with every record

A model field makes mixed-model mistakes detectable and re-embedding auditable.

Quick check: Why derive record IDs from document ID and chunk number?

  • Databases forbid other IDs
  • IDs must be random
  • It makes vectors smaller
  • Re-ingesting overwrites instead of duplicating
Answer

Re-ingesting overwrites instead of duplicating — Stable IDs make updates idempotent.