Lesson 25 / 25

Revision: Cheat Sheet and Self-Check

Review Spark's concepts, APIs and tuning habits from the whole course.

Cheat sheet

Architecture: driver plans; executors run tasks on partitions; action → job → stages (split at shuffles) → tasks. API: SparkSession; DataFrames (immutable) with select/filter/withColumn/groupBy/agg/join/window; prefer built-ins to UDFs; SQL = DataFrame API. Execution: transformations lazy, actions run; narrow vs wide; Exchange = shuffle; explain(); AQE. Partitions: ~100-200 MB; repartition (shuffle) vs coalesce (merge); fix skew (AQE, salting, broadcast). Performance: cache reused data, broadcast small joins, Parquet + partition pruning, avoid small files, tune few configs. Streaming: unbounded table, trigger, output mode, checkpoint, event-time windows + watermark. Production: object storage, YARN/Kubernetes/managed, function-style code + tests + data-quality checks, Spark UI symptoms, Delta/Iceberg, idempotent partition overwrite.

Questions interviewers ask

Be ready to explain: the difference between transformations and actions and lazy evaluation, what causes a shuffle and how to reduce it, narrow versus wide dependencies, repartition versus coalesce, how you would handle a skewed join, when to broadcast, how caching works, and how you would debug a slow Spark job from the Spark UI.

Quick check: A job calls `df.count()` twice on an expensive DataFrame with no caching. What happens?

  • Spark errors out
  • The second call is free
  • The whole lineage is computed twice
  • The result is cached automatically
Answer

The whole lineage is computed twice — Each action recomputes from the source unless the DataFrame is cached or persisted.

Quick check: A join between a 2 TB table and a 5 MB lookup table is slow. What is the best first idea?

  • Broadcast the 5 MB table
  • Collect the 2 TB table to the driver
  • Add more shuffle partitions only
  • Convert both to CSV
Answer

Broadcast the 5 MB table — Broadcasting the small side avoids shuffling the huge table.

Quick check: Which statement about coalesce and repartition is correct?

  • repartition never moves data
  • They are identical
  • coalesce can increase partitions with a full shuffle
  • coalesce only reduces partitions without a full shuffle; repartition shuffles to an exact count
Answer

coalesce only reduces partitions without a full shuffle; repartition shuffles to an exact count — coalesce is a cheap merge for shrinking; repartition redistributes all data evenly.