Lesson 2 / 25
Architecture: Driver, Executors, Jobs, Stages, Tasks
Name the runtime components and how an action becomes jobs, stages and tasks.
From your code to tasks
Your program runs as the driver, which holds the SparkSession and builds the plan. A cluster manager (Spark standalone, YARN, Kubernetes or others) provides resources, and executors are JVM processes on worker machines that run tasks and cache data. Each action (like count() or write) triggers a job. Spark splits a job into stages at shuffle boundaries (where data must move between machines), and each stage into tasks, one per partition. Tasks of a stage run in parallel; a stage begins only when the stages it depends on have finished.
The hierarchy
Read it top to bottom. Fewer, larger stages and well-sized partitions usually mean better performance.
Application (your program, one SparkSession)
└─ Job (one per action: count, collect, write, show...)
└─ Stage (a set of tasks with no shuffle between them)
└─ Task (one partition processed by one executor core)
Driver : plans, schedules, collects small results
Executors : run tasks, cache data, write output
Cluster manager : gives the application CPU and memorycollect() brings everything to the driver
collect() and toPandas() copy all rows into the driver's memory. On a large dataset this causes out-of-memory errors. Use show(), limit() or write results to storage instead.
Quick check: What triggers a Spark job?
- An action such as count() or write
- Creating a DataFrame variable
- Importing pyspark
- Typing a comment
Answer
An action such as count() or write — Transformations only build a plan; an action makes Spark actually execute it.