पाठ 18 / 25
Data Layout, छोटी Files और Configuration
छोटी-file समस्याएँ ठीक करें, file आकार चुनें और वे कुछ configs तय करें जो मायने रखते हैं।
असली धीमापन ज़्यादातर कहाँ छिपता है
छोटी files (कुछ KB की हज़ारों files) पढ़ना धीमा करती हैं क्योंकि हर file एक open call और एक task की क़ीमत लेती है। लिखने से पहले output partitions नियंत्रित करके (coalesce/repartition), समय-समय पर compact करके, या files को compact और प्रबंधित करने वाले table formats, जैसे Delta Lake या Apache Iceberg, उपयोग करके इनसे बचें। Columnar formats उपयोग करें और सिर्फ़ ज़रूरी columns चुनें। कुछ settings सबसे ज़्यादा मायने रखती हैं: spark.sql.shuffle.partitions (shuffle की चौड़ाई; AQE मदद करता है), executor memory और cores (आम नियम प्रति executor 4-5 cores, विशाल heaps से बचना), spark.sql.adaptive.enabled (चालू रखें), और serialization (RDD-भारी jobs के लिए Kryo)। जहाँ built-in function हो वहाँ Python UDFs से बचें, क्योंकि हर row JVM-Python सीमा पार करती है; कस्टम तर्क चाहिए तो pandas UDFs (vectorised) या native SQL expressions चुनें।
UDF बनाम built-in, चलाकर
मैंने यह Apache Spark 4.0.0 (PySpark, local mode, official Docker image में) पर चलाया। दोनों तरीक़े वही उत्तर, ASHA!, देते हैं, पर built-in upper/concat संस्करण JVM के भीतर चलता है और अनुकूलित हो सकता है, जबकि Python UDF Catalyst के लिए अपारदर्शी है और बड़े पैमाने पर धीमा।
from pyspark.sql.types import StringType
shout = F.udf(lambda s: s.upper() + "!", StringType())
print(orders.select(shout("customer").alias("u")).orderBy("u").collect()[0][0], orders.select(F.concat(F.upper("customer"), F.lit("!")).alias("b")).orderBy("b").collect()[0][0])
Output:
ASHA! ASHA!
कुछ उपयोगी configs (उदाहरण)
डिफ़ॉल्ट से शुरू करें, एक बार में एक चीज़ बदलें और Spark UI में मापें। मान cluster के आकार और डेटा पर निर्भर हैं।
spark = (SparkSession.builder
.config("spark.sql.shuffle.partitions", "400") # AQE can coalesce this
.config("spark.sql.adaptive.enabled", "true")
.config("spark.sql.autoBroadcastJoinThreshold", str(50 * 1024 * 1024))
.config("spark.executor.memory", "8g")
.config("spark.executor.cores", "4")
.getOrCreate())त्वरित जाँच: हज़ारों छोटी files से क्यों बचें?
- हर file open और task का बोझ जोड़ती है, जो reads धीमे करता है
- छोटी files अवैध हैं
- वे बहुत कम disk लेती हैं
- Spark उन्हें बिल्कुल नहीं पढ़ सकता
Answer
हर file open और task का बोझ जोड़ती है, जो reads धीमे करता है — फ़ाइलें छोटी हों तो प्रति-file बोझ हावी होता है, इसलिए उन्हें बड़ी files में compact करें।