Apache Spark: Big Data Processing with DataFrames

Learn Apache Spark with PySpark: architecture, DataFrames, SQL, joins, windows, lazy execution, shuffles, performance tuning, streaming and production deployment.

Start course →

Syllabus

Spark Basics

  1. What Spark Is
  2. Architecture: Driver, Executors, Jobs, Stages, Tasks
  3. Running Spark and Creating a SparkSession
  4. RDDs, DataFrames and Spark SQL

The DataFrame API

  1. Selecting, Filtering and Columns
  2. Aggregations and Null Handling
  3. Joins
  4. Window Functions

Spark SQL and Files

  1. Spark SQL and Temporary Views
  2. File Formats: CSV, JSON, Parquet
  3. Partitioned Output and Partition Pruning

How Spark Executes Your Code

  1. Transformations, Actions and Lazy Evaluation
  2. Narrow and Wide Transformations: the Shuffle
  3. Reading Plans with explain() and Adaptive Query Execution
  4. Partitions: repartition, coalesce and Skew

Performance Tuning

  1. Caching and Persistence
  2. Join Strategies and Broadcast Joins
  3. Data Layout, Small Files and Configuration

Structured Streaming

  1. The Structured Streaming Model
  2. Event-Time Windows and Watermarks

Production Spark

  1. Cluster Managers and Platforms
  2. Testing Spark Jobs
  3. Monitoring with the Spark UI and Lakehouse Tables

Putting It Together

  1. Case Study: A Daily Sales Pipeline
  2. Revision: Cheat Sheet and Self-Check