पाठ 23 / 25
Spark UI से Monitoring और Lakehouse Tables
Spark UI में धीमे jobs का निदान करें और भरोसा बढ़ाने वाले table formats उपयोग करें।
Job धीमा हो तो कहाँ देखें
Spark UI (चलते application के लिए port 4040, या ख़त्म हुए के लिए History Server) आपका मुख्य साधन है। Jobs और Stages दिखाते हैं कि समय कहाँ जाता है; stage में tasks की अवधियाँ तुलना करें: कुछ बहुत लंबे tasks का मतलब skew, विशाल shuffle read/write का मतलब महँगा shuffle, disk पर spill का मतलब प्रति task कम memory, और हज़ारों छोटे tasks का मतलब बहुत ज़्यादा partitions या छोटी files। SQL tab metrics के साथ query plan दिखाता है, Executors memory उपयोग और garbage-collection समय दिखाता है, और Storage cached डेटा। भरोसेमंद tables के लिए Delta Lake और Apache Iceberg object storage में Parquet files के ऊपर ACID transactions, schema evolution, time travel और कुशल upserts (MERGE) जोड़ते हैं, जिससे idempotent pipelines और सुरक्षित दोबारा processing कहीं आसान होती है।
लक्षण पढ़ना (cheat sheet)
Spark UI में दिखने वाले लक्षण को संभावित कारण और पहले सुधार से मिलाएँ।
Symptom in the UI Likely cause First thing to try
one task >> others in a stage data skew AQE skew join, salting, broadcast
huge shuffle read/write wide ops on too much data filter/select earlier, broadcast, fewer shuffles
spill (memory/disk) high tasks too big for memory more partitions, more executor memory
thousands of 10 KB tasks small files / over-partitioned compact files, coalesce, AQE
long GC time heap too big / many objects smaller executors, avoid Python UDFs, use ParquetDelta Lake MERGE से upsert (उदाहरण)
एक ही दिन के लिए इसे दोबारा चलाने पर वही table मिलती है, जिससे backfills सुरक्षित होते हैं। Delta Lake package चाहिए।
MERGE INTO analytics.daily_orders AS t
USING staging.orders_today AS s
ON t.order_id = s.order_id
WHEN MATCHED THEN UPDATE SET *
WHEN NOT MATCHED THEN INSERT *;त्वरित जाँच: Spark UI में एक stage का एक task बाक़ी से कहीं ज़्यादा देर चलता है। यह क्या सुझाता है?
- Data skew
- तेज़ network
- Job पूरा हो गया
- बहुत कम tasks हैं
Answer
Data skew — पिछड़ता task आम तौर पर ऐसी गर्म key रखता है जिसका डेटा बाक़ियों से कहीं ज़्यादा है।