# Data Layout, छोटी Files और Configuration — Apache Spark: DataFrames से Big Data Processing

Source: https://www.geekswithgeeks.com/hi/spark/perf-layout-config

> छोटी-file समस्याएँ ठीक करें, file आकार चुनें और वे कुछ configs तय करें जो मायने रखते हैं।

## असली धीमापन ज़्यादातर कहाँ छिपता है

**छोटी files** (कुछ KB की हज़ारों files) पढ़ना धीमा करती हैं क्योंकि हर file एक open call और एक task की क़ीमत लेती है। लिखने से पहले output partitions नियंत्रित करके (`coalesce`/`repartition`), समय-समय पर compact करके, या files को compact और प्रबंधित करने वाले table formats, जैसे **Delta Lake** या **Apache Iceberg**, उपयोग करके इनसे बचें। **Columnar formats** उपयोग करें और सिर्फ़ ज़रूरी columns चुनें। कुछ settings सबसे ज़्यादा मायने रखती हैं: `spark.sql.shuffle.partitions` (shuffle की चौड़ाई; AQE मदद करता है), executor `memory` और `cores` (आम नियम प्रति executor 4-5 cores, विशाल heaps से बचना), `spark.sql.adaptive.enabled` (चालू रखें), और serialization (RDD-भारी jobs के लिए Kryo)। जहाँ built-in function हो वहाँ Python UDFs से बचें, क्योंकि हर row JVM-Python सीमा पार करती है; कस्टम तर्क चाहिए तो **pandas UDFs** (vectorised) या native SQL expressions चुनें।

## UDF बनाम built-in, चलाकर

मैंने यह Apache Spark 4.0.0 (PySpark, local mode, official Docker image में) पर चलाया। दोनों तरीक़े वही उत्तर, `ASHA!`, देते हैं, पर built-in `upper`/`concat` संस्करण JVM के भीतर चलता है और अनुकूलित हो सकता है, जबकि Python UDF Catalyst के लिए अपारदर्शी है और बड़े पैमाने पर धीमा।

```python
from pyspark.sql.types import StringType
shout = F.udf(lambda s: s.upper() + "!", StringType())
print(orders.select(shout("customer").alias("u")).orderBy("u").collect()[0][0], orders.select(F.concat(F.upper("customer"), F.lit("!")).alias("b")).orderBy("b").collect()[0][0])
```

Output:

```
ASHA! ASHA!
```

## कुछ उपयोगी configs (उदाहरण)

डिफ़ॉल्ट से शुरू करें, एक बार में एक चीज़ बदलें और Spark UI में मापें। मान cluster के आकार और डेटा पर निर्भर हैं।

```python
spark = (SparkSession.builder
    .config("spark.sql.shuffle.partitions", "400")          # AQE can coalesce this
    .config("spark.sql.adaptive.enabled", "true")
    .config("spark.sql.autoBroadcastJoinThreshold", str(50 * 1024 * 1024))
    .config("spark.executor.memory", "8g")
    .config("spark.executor.cores", "4")
    .getOrCreate())
```

**Quiz:** हज़ारों छोटी files से क्यों बचें?

- [x] हर file open और task का बोझ जोड़ती है, जो reads धीमे करता है
- [ ] छोटी files अवैध हैं
- [ ] वे बहुत कम disk लेती हैं
- [ ] Spark उन्हें बिल्कुल नहीं पढ़ सकता

*Answer:* हर file open और task का बोझ जोड़ती है, जो reads धीमे करता है. फ़ाइलें छोटी हों तो प्रति-file बोझ हावी होता है, इसलिए उन्हें बड़ी files में compact करें।
