# Chroma: A Developer-Friendly Embedded Store — Vector Databases

Source: https://www.geekswithgeeks.com/en/vector-databases/e-chroma

> Use a lightweight engine for prototypes, notebooks and small apps.

## Quick to start, fewer knobs

**Chroma** is an open-source embedding database aimed at developer productivity. You create a **collection**, `add` records (ids, embeddings or documents, metadata) and `query` with embeddings or text, optionally with a `where` metadata filter. It can run **in-process** (in memory or persisted to a folder) or as a **server**. If you give it documents without embeddings it calls an **embedding function** for you, which may download or call a model; in the example below we pass our own vectors so nothing is downloaded. Embedded stores like this are excellent for prototypes, tests and small single-user apps; for large multi-user production workloads compare it with the other families on scale, filtering, replication and operations.

## Same ideas, different APIs

Most engines expose collections, upserts, filtered queries and deletes; compare them on your workload, not on a feature list.

![Three questions: API, features, evidence.](assets/figures/vector-databases/section-5-map.svg) — Figure 5.1 — API, features and evidence.

## Chroma: add, query, filter, delete, run

I ran this Python in a virtual environment with qdrant-client 1.19.1 (local in-memory mode) and chromadb 1.5.9 (in-memory client). Using cosine distance, the nearest records to `[1,0,0]` are a, d, b with small distances. A `where` filter for tenant `acme` and year 2025 leaves a and c. After deleting d the count is 3.

```python
import chromadb

client = chromadb.EphemeralClient()                       # in-memory; we pass our own vectors, so no model download
col = client.create_collection("docs", metadata={"hnsw:space": "cosine"})
col.add(
    ids=["a", "b", "c", "d"],
    embeddings=[[0.9, 0.1, 0.0], [0.8, 0.2, 0.1], [0.1, 0.9, 0.1], [0.9, 0.0, 0.2]],
    documents=["leave policy 2025", "old leave policy", "travel and hotels", "beta confidential memo"],
    metadatas=[{"tenant": "acme", "year": 2025}, {"tenant": "acme", "year": 2022},
               {"tenant": "acme", "year": 2025}, {"tenant": "beta", "year": 2025}],
)
r = col.query(query_embeddings=[[1, 0, 0]], n_results=3)
print("no filter :", r["ids"][0], [round(d, 3) for d in r["distances"][0]])
r = col.query(query_embeddings=[[1, 0, 0]], n_results=3, where={"$and": [{"tenant": "acme"}, {"year": 2025}]})
print("with where:", r["ids"][0], r["documents"][0])
col.delete(ids=["d"])
print("count after delete:", col.count())

```

Output:

```
no filter : ['a', 'd', 'b'] [0.006, 0.024, 0.037]
with where: ['a', 'c'] ['leave policy 2025', 'travel and hotels']
count after delete: 3
```

## Pass your own embeddings in production

Controlling the embedding model yourself keeps vectors reproducible and avoids surprise model downloads.

**Quiz:** What does Chroma do if you add documents without embeddings?

- [ ] Stores them as images
- [ ] Refuses all inserts
- [x] Calls a configured embedding function to create them
- [ ] Deletes the collection

*Answer:* Calls a configured embedding function to create them. Convenient, but it may download or call a model; pass your own vectors for control.
