# Monitoring with the Spark UI and Lakehouse Tables — Apache Spark: Big Data Processing with DataFrames

Source: https://www.geekswithgeeks.com/en/spark/prod-monitoring-tables

> Diagnose slow jobs in the Spark UI and use table formats that add reliability.

## Where to look when a job is slow

The **Spark UI** (port 4040 for a running application, or the History Server for finished ones) is your main tool. **Jobs and Stages** show where time goes; in a stage, compare task durations: a few very long tasks mean **skew**, huge **shuffle read/write** means an expensive shuffle, **spill** to disk means too little memory per task, and thousands of tiny tasks mean too many partitions or small files. The **SQL** tab shows the query plan with metrics, **Executors** shows memory use and garbage-collection time, and **Storage** shows cached data. For reliable tables, **Delta Lake** and **Apache Iceberg** add ACID transactions, schema evolution, time travel and efficient upserts (`MERGE`) on top of Parquet files in object storage, which makes idempotent pipelines and safe reprocessing much easier.

## Reading symptoms (a cheat sheet)

Match the symptom you see in the Spark UI to the likely cause and first fix.

```text
Symptom in the UI                          Likely cause                 First thing to try
one task >> others in a stage                data skew                    AQE skew join, salting, broadcast
huge shuffle read/write                      wide ops on too much data    filter/select earlier, broadcast, fewer shuffles
spill (memory/disk) high                     tasks too big for memory     more partitions, more executor memory
thousands of 10 KB tasks                     small files / over-partitioned compact files, coalesce, AQE
long GC time                                 heap too big / many objects  smaller executors, avoid Python UDFs, use Parquet
```

## An upsert with Delta Lake MERGE (illustrative)

Re-running this for the same day gives the same table, which is what makes backfills safe. Requires the Delta Lake package.

```sql
MERGE INTO analytics.daily_orders AS t
USING staging.orders_today AS s
  ON t.order_id = s.order_id
WHEN MATCHED THEN UPDATE SET *
WHEN NOT MATCHED THEN INSERT *;
```

**Quiz:** In the Spark UI a stage has one task running far longer than the rest. What does it suggest?

- [x] Data skew
- [ ] A fast network
- [ ] The job is finished
- [ ] Too few tasks exist

*Answer:* Data skew. A straggler task usually holds a hot key with far more data than the others.
