Lesson 25 / 25
Revision: Cheat Sheet and Self-Check
Review Spark's concepts, APIs and tuning habits from the whole course.
Cheat sheet
Architecture: driver plans; executors run tasks on partitions; action → job → stages (split at shuffles) → tasks. API: SparkSession; DataFrames (immutable) with select/filter/withColumn/groupBy/agg/join/window; prefer built-ins to UDFs; SQL = DataFrame API. Execution: transformations lazy, actions run; narrow vs wide; Exchange = shuffle; explain(); AQE. Partitions: ~100-200 MB; repartition (shuffle) vs coalesce (merge); fix skew (AQE, salting, broadcast). Performance: cache reused data, broadcast small joins, Parquet + partition pruning, avoid small files, tune few configs. Streaming: unbounded table, trigger, output mode, checkpoint, event-time windows + watermark. Production: object storage, YARN/Kubernetes/managed, function-style code + tests + data-quality checks, Spark UI symptoms, Delta/Iceberg, idempotent partition overwrite.
Questions interviewers ask
Be ready to explain: the difference between transformations and actions and lazy evaluation, what causes a shuffle and how to reduce it, narrow versus wide dependencies, repartition versus coalesce, how you would handle a skewed join, when to broadcast, how caching works, and how you would debug a slow Spark job from the Spark UI.
Quick check: A job calls `df.count()` twice on an expensive DataFrame with no caching. What happens?
- Spark errors out
- The second call is free
- The whole lineage is computed twice
- The result is cached automatically
Answer
The whole lineage is computed twice — Each action recomputes from the source unless the DataFrame is cached or persisted.
Quick check: A join between a 2 TB table and a 5 MB lookup table is slow. What is the best first idea?
- Broadcast the 5 MB table
- Collect the 2 TB table to the driver
- Add more shuffle partitions only
- Convert both to CSV
Answer
Broadcast the 5 MB table — Broadcasting the small side avoids shuffling the huge table.
Quick check: Which statement about coalesce and repartition is correct?
- repartition never moves data
- They are identical
- coalesce can increase partitions with a full shuffle
- coalesce only reduces partitions without a full shuffle; repartition shuffles to an exact count
Answer
coalesce only reduces partitions without a full shuffle; repartition shuffles to an exact count — coalesce is a cheap merge for shrinking; repartition redistributes all data evenly.