पाठ 21 / 28
Sharding: डेटा को मशीनों में बाँटना
Hash sharding समझें और shard संख्या बदलना महँगा क्यों है।
हर shard एक हिस्सा रखता है; queries फैलती हैं
जब एक मशीन सारे vectors नहीं रख या परोस सकती, तब डेटा shard किया जाता है: हर record एक shard को दिया जाता है, आम तौर पर उसकी ID (या tenant key) के hash से। Query हर shard को भेजी जाती है (या उन shards को जिनमें मेल हो सकते हैं), हर एक अपना स्थानीय top-k लौटाता है, और coordinator उन्हें वैश्विक top-k में मिलाता है। ज़्यादा shards का अर्थ ज़्यादा memory और समानांतरता, पर हर query ज़्यादा मशीनों को छूती है, इसलिए tail latency सबसे धीमे shard से तय होती है। Shard key को पहुँच पैटर्न से मिलाएँ (tenant से रखने पर एक tenant कुछ shards पर रहता है और tenant queries सस्ती होती हैं)। Hash-आधारित ढाँचे का resharding अधिकांश records हिलाता है, इसलिए shards की संख्या पहले से योजना में रखें (या consistent hashing या कई छोटे logical shards वाले engines उपयोग करें जो इकाई के रूप में हिलते हैं)।
Hash sharding और बदलाव की क़ीमत, चलाकर
मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। 10,000 IDs का hash 4 shards में से हर एक को लगभग 2,500 records और 8 shards में हर एक को लगभग 1,250 देता है, समान फैलाव। पर hash % N के साथ 4 से 5 shards पर जाने पर 10,000 में से 7,921 records, लगभग 80%, हिलते हैं, इसीलिए resharding महँगा है।
import hashlib
from collections import Counter
def shard_for(doc_id, n_shards):
return int(hashlib.md5(doc_id.encode()).hexdigest(), 16) % n_shards
ids = [f"doc-{i}" for i in range(10000)]
for n in (4, 8):
print(n, "shards ->", sorted(Counter(shard_for(i, n) for i in ids).values()))
moved = sum(shard_for(i, 4) != shard_for(i, 5) for i in ids)
print("going from 4 to 5 shards with hash % N moves", moved, "of", len(ids), "documents")
Output:
4 shards -> [2462, 2468, 2519, 2551] 8 shards -> [1204, 1214, 1247, 1254, 1258, 1258, 1272, 1293] going from 4 to 5 shards with hash % N moves 7921 of 10000 documents
वृद्धि को ध्यान में रखकर shard संख्या चुनें
Hash-आधारित resharding अधिकांश डेटा हिलाता है, इसलिए ज़रूरत पड़ने से पहले बढ़ने की जगह रखें।
त्वरित जाँच: Sharded vector database में query का उत्तर कैसे मिलता है?
- सिर्फ़ पहला shard खोजा जाता है
- हर shard अपना स्थानीय top-k लौटाता है और coordinator उन्हें मिलाता है
- Client disks खोजता है
- Shards लिखते समय मिलाए जाते हैं
Answer
हर shard अपना स्थानीय top-k लौटाता है और coordinator उन्हें मिलाता है — Fan-out और merge वैश्विक निकटतम पड़ोसी देता है।