Lesson 23 / 25

Monitoring with the Spark UI and Lakehouse Tables

Diagnose slow jobs in the Spark UI and use table formats that add reliability.

Where to look when a job is slow

The Spark UI (port 4040 for a running application, or the History Server for finished ones) is your main tool. Jobs and Stages show where time goes; in a stage, compare task durations: a few very long tasks mean skew, huge shuffle read/write means an expensive shuffle, spill to disk means too little memory per task, and thousands of tiny tasks mean too many partitions or small files. The SQL tab shows the query plan with metrics, Executors shows memory use and garbage-collection time, and Storage shows cached data. For reliable tables, Delta Lake and Apache Iceberg add ACID transactions, schema evolution, time travel and efficient upserts (MERGE) on top of Parquet files in object storage, which makes idempotent pipelines and safe reprocessing much easier.

Reading symptoms (a cheat sheet)

Match the symptom you see in the Spark UI to the likely cause and first fix.

Symptom in the UI                          Likely cause                 First thing to try
one task >> others in a stage                data skew                    AQE skew join, salting, broadcast
huge shuffle read/write                      wide ops on too much data    filter/select earlier, broadcast, fewer shuffles
spill (memory/disk) high                     tasks too big for memory     more partitions, more executor memory
thousands of 10 KB tasks                     small files / over-partitioned compact files, coalesce, AQE
long GC time                                 heap too big / many objects  smaller executors, avoid Python UDFs, use Parquet

An upsert with Delta Lake MERGE (illustrative)

Re-running this for the same day gives the same table, which is what makes backfills safe. Requires the Delta Lake package.

MERGE INTO analytics.daily_orders AS t
USING staging.orders_today AS s
  ON t.order_id = s.order_id
WHEN MATCHED THEN UPDATE SET *
WHEN NOT MATCHED THEN INSERT *;

Quick check: In the Spark UI a stage has one task running far longer than the rest. What does it suggest?

  • Data skew
  • A fast network
  • The job is finished
  • Too few tasks exist
Answer

Data skew — A straggler task usually holds a hot key with far more data than the others.