# Revision: Cheat Sheet and Self-Check — Apache Spark: Big Data Processing with DataFrames

Source: https://www.geekswithgeeks.com/en/spark/wrap-revision

> Review Spark's concepts, APIs and tuning habits from the whole course.

## Cheat sheet

**Architecture**: driver plans; executors run tasks on partitions; action → job → stages (split at shuffles) → tasks. **API**: SparkSession; DataFrames (immutable) with select/filter/withColumn/groupBy/agg/join/window; prefer built-ins to UDFs; SQL = DataFrame API. **Execution**: transformations lazy, actions run; narrow vs wide; `Exchange` = shuffle; `explain()`; AQE. **Partitions**: ~100-200 MB; `repartition` (shuffle) vs `coalesce` (merge); fix skew (AQE, salting, broadcast). **Performance**: cache reused data, broadcast small joins, Parquet + partition pruning, avoid small files, tune few configs. **Streaming**: unbounded table, trigger, output mode, checkpoint, event-time windows + watermark. **Production**: object storage, YARN/Kubernetes/managed, function-style code + tests + data-quality checks, Spark UI symptoms, Delta/Iceberg, idempotent partition overwrite.

## Questions interviewers ask

Be ready to explain: the difference between transformations and actions and lazy evaluation, what causes a shuffle and how to reduce it, narrow versus wide dependencies, repartition versus coalesce, how you would handle a skewed join, when to broadcast, how caching works, and how you would debug a slow Spark job from the Spark UI.

**Quiz:** A job calls `df.count()` twice on an expensive DataFrame with no caching. What happens?

- [ ] Spark errors out
- [ ] The second call is free
- [x] The whole lineage is computed twice
- [ ] The result is cached automatically

*Answer:* The whole lineage is computed twice. Each action recomputes from the source unless the DataFrame is cached or persisted.

**Quiz:** A join between a 2 TB table and a 5 MB lookup table is slow. What is the best first idea?

- [x] Broadcast the 5 MB table
- [ ] Collect the 2 TB table to the driver
- [ ] Add more shuffle partitions only
- [ ] Convert both to CSV

*Answer:* Broadcast the 5 MB table. Broadcasting the small side avoids shuffling the huge table.

**Quiz:** Which statement about coalesce and repartition is correct?

- [ ] repartition never moves data
- [ ] They are identical
- [ ] coalesce can increase partitions with a full shuffle
- [x] coalesce only reduces partitions without a full shuffle; repartition shuffles to an exact count

*Answer:* coalesce only reduces partitions without a full shuffle; repartition shuffles to an exact count. coalesce is a cheap merge for shrinking; repartition redistributes all data evenly.
