Lesson 23 / 25
Monitoring with the Spark UI and Lakehouse Tables
Diagnose slow jobs in the Spark UI and use table formats that add reliability.
Where to look when a job is slow
The Spark UI (port 4040 for a running application, or the History Server for finished ones) is your main tool. Jobs and Stages show where time goes; in a stage, compare task durations: a few very long tasks mean skew, huge shuffle read/write means an expensive shuffle, spill to disk means too little memory per task, and thousands of tiny tasks mean too many partitions or small files. The SQL tab shows the query plan with metrics, Executors shows memory use and garbage-collection time, and Storage shows cached data. For reliable tables, Delta Lake and Apache Iceberg add ACID transactions, schema evolution, time travel and efficient upserts (MERGE) on top of Parquet files in object storage, which makes idempotent pipelines and safe reprocessing much easier.
Reading symptoms (a cheat sheet)
Match the symptom you see in the Spark UI to the likely cause and first fix.
Symptom in the UI Likely cause First thing to try
one task >> others in a stage data skew AQE skew join, salting, broadcast
huge shuffle read/write wide ops on too much data filter/select earlier, broadcast, fewer shuffles
spill (memory/disk) high tasks too big for memory more partitions, more executor memory
thousands of 10 KB tasks small files / over-partitioned compact files, coalesce, AQE
long GC time heap too big / many objects smaller executors, avoid Python UDFs, use ParquetAn upsert with Delta Lake MERGE (illustrative)
Re-running this for the same day gives the same table, which is what makes backfills safe. Requires the Delta Lake package.
MERGE INTO analytics.daily_orders AS t
USING staging.orders_today AS s
ON t.order_id = s.order_id
WHEN MATCHED THEN UPDATE SET *
WHEN NOT MATCHED THEN INSERT *;Quick check: In the Spark UI a stage has one task running far longer than the rest. What does it suggest?
- Data skew
- A fast network
- The job is finished
- Too few tasks exist
Answer
Data skew — A straggler task usually holds a hot key with far more data than the others.