Apache Spark: Big Data Processing with DataFrames
Learn Apache Spark with PySpark: architecture, DataFrames, SQL, joins, windows, lazy execution, shuffles, performance tuning, streaming and production deployment.
Start course →
Syllabus
Spark Basics
- What Spark Is
- Architecture: Driver, Executors, Jobs, Stages, Tasks
- Running Spark and Creating a SparkSession
- RDDs, DataFrames and Spark SQL
The DataFrame API
- Selecting, Filtering and Columns
- Aggregations and Null Handling
- Joins
- Window Functions
Spark SQL and Files
- Spark SQL and Temporary Views
- File Formats: CSV, JSON, Parquet
- Partitioned Output and Partition Pruning
How Spark Executes Your Code
- Transformations, Actions and Lazy Evaluation
- Narrow and Wide Transformations: the Shuffle
- Reading Plans with explain() and Adaptive Query Execution
- Partitions: repartition, coalesce and Skew
Performance Tuning
- Caching and Persistence
- Join Strategies and Broadcast Joins
- Data Layout, Small Files and Configuration
Structured Streaming
- The Structured Streaming Model
- Event-Time Windows and Watermarks
Production Spark
- Cluster Managers and Platforms
- Testing Spark Jobs
- Monitoring with the Spark UI and Lakehouse Tables
Putting It Together
- Case Study: A Daily Sales Pipeline
- Revision: Cheat Sheet and Self-Check