# Sharding: डेटा को मशीनों में बाँटना — Vector Databases

Source: https://www.geekswithgeeks.com/hi/vector-databases/o-shard

> Hash sharding समझें और shard संख्या बदलना महँगा क्यों है।

## हर shard एक हिस्सा रखता है; queries फैलती हैं

जब एक मशीन सारे vectors नहीं रख या परोस सकती, तब डेटा **shard** किया जाता है: हर record एक shard को दिया जाता है, आम तौर पर उसकी ID (या tenant key) के hash से। Query **हर shard को भेजी जाती है** (या उन shards को जिनमें मेल हो सकते हैं), हर एक अपना स्थानीय top-k लौटाता है, और coordinator उन्हें वैश्विक top-k में **मिलाता** है। ज़्यादा shards का अर्थ ज़्यादा memory और समानांतरता, पर हर query ज़्यादा मशीनों को छूती है, इसलिए **tail latency** सबसे धीमे shard से तय होती है। **Shard key** को पहुँच पैटर्न से मिलाएँ (tenant से रखने पर एक tenant कुछ shards पर रहता है और tenant queries सस्ती होती हैं)। Hash-आधारित ढाँचे का resharding अधिकांश records हिलाता है, इसलिए **shards की संख्या पहले से** योजना में रखें (या consistent hashing या कई छोटे logical shards वाले engines उपयोग करें जो इकाई के रूप में हिलते हैं)।

## Hash sharding और बदलाव की क़ीमत, चलाकर

मैंने यह सादा-Python (सिर्फ़ standard library) उदाहरण चलाया। 10,000 IDs का hash 4 shards में से हर एक को लगभग 2,500 records और 8 shards में हर एक को लगभग 1,250 देता है, समान फैलाव। पर `hash % N` के साथ 4 से 5 shards पर जाने पर 10,000 में से 7,921 records, लगभग 80%, हिलते हैं, इसीलिए resharding महँगा है।

```python
import hashlib
from collections import Counter

def shard_for(doc_id, n_shards):
    return int(hashlib.md5(doc_id.encode()).hexdigest(), 16) % n_shards

ids = [f"doc-{i}" for i in range(10000)]
for n in (4, 8):
    print(n, "shards ->", sorted(Counter(shard_for(i, n) for i in ids).values()))

moved = sum(shard_for(i, 4) != shard_for(i, 5) for i in ids)
print("going from 4 to 5 shards with hash % N moves", moved, "of", len(ids), "documents")

```

Output:

```
4 shards -> [2462, 2468, 2519, 2551]
8 shards -> [1204, 1214, 1247, 1254, 1258, 1258, 1272, 1293]
going from 4 to 5 shards with hash % N moves 7921 of 10000 documents
```

## वृद्धि को ध्यान में रखकर shard संख्या चुनें

Hash-आधारित resharding अधिकांश डेटा हिलाता है, इसलिए ज़रूरत पड़ने से पहले बढ़ने की जगह रखें।

**Quiz:** Sharded vector database में query का उत्तर कैसे मिलता है?

- [ ] सिर्फ़ पहला shard खोजा जाता है
- [x] हर shard अपना स्थानीय top-k लौटाता है और coordinator उन्हें मिलाता है
- [ ] Client disks खोजता है
- [ ] Shards लिखते समय मिलाए जाते हैं

*Answer:* हर shard अपना स्थानीय top-k लौटाता है और coordinator उन्हें मिलाता है. Fan-out और merge वैश्विक निकटतम पड़ोसी देता है।
