# Spark UI से Monitoring और Lakehouse Tables — Apache Spark: DataFrames से Big Data Processing

Source: https://www.geekswithgeeks.com/hi/spark/prod-monitoring-tables

> Spark UI में धीमे jobs का निदान करें और भरोसा बढ़ाने वाले table formats उपयोग करें।

## Job धीमा हो तो कहाँ देखें

**Spark UI** (चलते application के लिए port 4040, या ख़त्म हुए के लिए History Server) आपका मुख्य साधन है। **Jobs और Stages** दिखाते हैं कि समय कहाँ जाता है; stage में tasks की अवधियाँ तुलना करें: कुछ बहुत लंबे tasks का मतलब **skew**, विशाल **shuffle read/write** का मतलब महँगा shuffle, disk पर **spill** का मतलब प्रति task कम memory, और हज़ारों छोटे tasks का मतलब बहुत ज़्यादा partitions या छोटी files। **SQL** tab metrics के साथ query plan दिखाता है, **Executors** memory उपयोग और garbage-collection समय दिखाता है, और **Storage** cached डेटा। भरोसेमंद tables के लिए **Delta Lake** और **Apache Iceberg** object storage में Parquet files के ऊपर ACID transactions, schema evolution, time travel और कुशल upserts (`MERGE`) जोड़ते हैं, जिससे idempotent pipelines और सुरक्षित दोबारा processing कहीं आसान होती है।

## लक्षण पढ़ना (cheat sheet)

Spark UI में दिखने वाले लक्षण को संभावित कारण और पहले सुधार से मिलाएँ।

```text
Symptom in the UI                          Likely cause                 First thing to try
one task >> others in a stage                data skew                    AQE skew join, salting, broadcast
huge shuffle read/write                      wide ops on too much data    filter/select earlier, broadcast, fewer shuffles
spill (memory/disk) high                     tasks too big for memory     more partitions, more executor memory
thousands of 10 KB tasks                     small files / over-partitioned compact files, coalesce, AQE
long GC time                                 heap too big / many objects  smaller executors, avoid Python UDFs, use Parquet
```

## Delta Lake MERGE से upsert (उदाहरण)

एक ही दिन के लिए इसे दोबारा चलाने पर वही table मिलती है, जिससे backfills सुरक्षित होते हैं। Delta Lake package चाहिए।

```sql
MERGE INTO analytics.daily_orders AS t
USING staging.orders_today AS s
  ON t.order_id = s.order_id
WHEN MATCHED THEN UPDATE SET *
WHEN NOT MATCHED THEN INSERT *;
```

**Quiz:** Spark UI में एक stage का एक task बाक़ी से कहीं ज़्यादा देर चलता है। यह क्या सुझाता है?

- [x] Data skew
- [ ] तेज़ network
- [ ] Job पूरा हो गया
- [ ] बहुत कम tasks हैं

*Answer:* Data skew. पिछड़ता task आम तौर पर ऐसी गर्म key रखता है जिसका डेटा बाक़ियों से कहीं ज़्यादा है।
