पाठ 17 / 28
Chroma: Developer-अनुकूल Embedded Store
Prototypes, notebooks और छोटे apps के लिए हल्का engine उपयोग करें।
शुरू करना जल्दी, कम knobs
Chroma developer उत्पादकता को लक्ष्य करने वाला open-source embedding database है। आप collection बनाते हैं, records (ids, embeddings या documents, metadata) add करते हैं और embeddings या पाठ से query करते हैं, वैकल्पिक रूप से where metadata filter के साथ। यह in-process (memory में या folder में सहेजकर) या server के रूप में चल सकता है। Documents को embeddings के बिना दें तो यह आपके लिए embedding function बुलाता है, जो मॉडल डाउनलोड या call कर सकता है; नीचे के उदाहरण में हम अपने vectors देते हैं ताकि कुछ डाउनलोड न हो। ऐसे embedded stores prototypes, tests और छोटे एकल-user apps के लिए उत्कृष्ट हैं; बड़े बहु-user production कार्यभार के लिए scale, filtering, replication और संचालन पर उसकी अन्य परिवारों से तुलना करें।
वही विचार, अलग APIs
अधिकांश engines collections, upserts, filtered queries और deletes देते हैं; उनकी तुलना feature सूची पर नहीं, अपने कार्यभार पर करें।
Chroma: add, query, filter, delete, चलाकर
मैंने यह Python virtual environment में qdrant-client 1.19.1 (स्थानीय in-memory mode) और chromadb 1.5.9 (in-memory client) के साथ चलाया। Cosine दूरी के साथ [1,0,0] के निकटतम records a, d, b हैं, छोटी दूरियों के साथ। Tenant acme और वर्ष 2025 का where filter a और c छोड़ता है। d हटाने के बाद गिनती 3 है।
import chromadb
client = chromadb.EphemeralClient() # in-memory; we pass our own vectors, so no model download
col = client.create_collection("docs", metadata={"hnsw:space": "cosine"})
col.add(
ids=["a", "b", "c", "d"],
embeddings=[[0.9, 0.1, 0.0], [0.8, 0.2, 0.1], [0.1, 0.9, 0.1], [0.9, 0.0, 0.2]],
documents=["leave policy 2025", "old leave policy", "travel and hotels", "beta confidential memo"],
metadatas=[{"tenant": "acme", "year": 2025}, {"tenant": "acme", "year": 2022},
{"tenant": "acme", "year": 2025}, {"tenant": "beta", "year": 2025}],
)
r = col.query(query_embeddings=[[1, 0, 0]], n_results=3)
print("no filter :", r["ids"][0], [round(d, 3) for d in r["distances"][0]])
r = col.query(query_embeddings=[[1, 0, 0]], n_results=3, where={"$and": [{"tenant": "acme"}, {"year": 2025}]})
print("with where:", r["ids"][0], r["documents"][0])
col.delete(ids=["d"])
print("count after delete:", col.count())
Output:
no filter : ['a', 'd', 'b'] [0.006, 0.024, 0.037] with where: ['a', 'c'] ['leave policy 2025', 'travel and hotels'] count after delete: 3
Production में अपने embeddings दें
Embedding model ख़ुद नियंत्रित करने से vectors दोहराने योग्य रहते हैं और अनपेक्षित मॉडल डाउनलोड से बचाव होता है।
त्वरित जाँच: दस्तावेज़ बिना embeddings के add करने पर Chroma क्या करता है?
- उन्हें images के रूप में रखता है
- सारे inserts अस्वीकार करता है
- उन्हें बनाने के लिए configured embedding function बुलाता है
- Collection हटाता है
Answer
उन्हें बनाने के लिए configured embedding function बुलाता है — सुविधाजनक, पर मॉडल डाउनलोड या call कर सकता है; नियंत्रण के लिए अपने vectors दें।