# Architecture: Driver, Executors, Jobs, Stages, Tasks — Apache Spark: Big Data Processing with DataFrames

Source: https://www.geekswithgeeks.com/en/spark/sp-architecture

> Name the runtime components and how an action becomes jobs, stages and tasks.

## From your code to tasks

Your program runs as the **driver**, which holds the `SparkSession` and builds the plan. A **cluster manager** (Spark standalone, YARN, Kubernetes or others) provides resources, and **executors** are JVM processes on worker machines that run tasks and cache data. Each **action** (like `count()` or `write`) triggers a **job**. Spark splits a job into **stages** at shuffle boundaries (where data must move between machines), and each stage into **tasks**, one per partition. Tasks of a stage run in parallel; a stage begins only when the stages it depends on have finished.

## The hierarchy

Read it top to bottom. Fewer, larger stages and well-sized partitions usually mean better performance.

```text
Application (your program, one SparkSession)
  └─ Job        (one per action: count, collect, write, show...)
       └─ Stage    (a set of tasks with no shuffle between them)
            └─ Task   (one partition processed by one executor core)

Driver    : plans, schedules, collects small results
Executors : run tasks, cache data, write output
Cluster manager : gives the application CPU and memory
```

## collect() brings everything to the driver

`collect()` and `toPandas()` copy all rows into the driver's memory. On a large dataset this causes out-of-memory errors. Use `show()`, `limit()` or write results to storage instead.

**Quiz:** What triggers a Spark job?

- [x] An action such as count() or write
- [ ] Creating a DataFrame variable
- [ ] Importing pyspark
- [ ] Typing a comment

*Answer:* An action such as count() or write. Transformations only build a plan; an action makes Spark actually execute it.
