# Architecture: Driver, Executors, Jobs, Stages, Tasks — Apache Spark: DataFrames से Big Data Processing

Source: https://www.geekswithgeeks.com/hi/spark/sp-architecture

> Runtime components के नाम बताएँ और जानें कि action कैसे jobs, stages और tasks बनता है।

## आपके कोड से tasks तक

आपका प्रोग्राम **driver** के रूप में चलता है, जो `SparkSession` रखता है और योजना बनाता है। **Cluster manager** (Spark standalone, YARN, Kubernetes या अन्य) संसाधन देता है, और **executors** worker मशीनों पर JVM processes हैं जो tasks चलाते और डेटा cache करते हैं। हर **action** (जैसे `count()` या `write`) एक **job** शुरू करता है। Spark job को shuffle सीमाओं पर (जहाँ डेटा मशीनों के बीच जाना होता है) **stages** में बाँटता है, और हर stage को **tasks** में, प्रति partition एक। एक stage के tasks समानांतर चलते हैं; stage तभी शुरू होता है जब उस पर निर्भर stages पूरे हो चुके हों।

## पदानुक्रम

इसे ऊपर से नीचे पढ़ें। कम, बड़े stages और सही आकार के partitions का आम तौर पर मतलब बेहतर performance है।

```text
Application (your program, one SparkSession)
  └─ Job        (one per action: count, collect, write, show...)
       └─ Stage    (a set of tasks with no shuffle between them)
            └─ Task   (one partition processed by one executor core)

Driver    : plans, schedules, collects small results
Executors : run tasks, cache data, write output
Cluster manager : gives the application CPU and memory
```

## collect() सब कुछ driver पर लाता है

`collect()` और `toPandas()` सारी rows driver की memory में कॉपी करते हैं। बड़े dataset पर इससे out-of-memory errors आते हैं। इसके बजाय `show()`, `limit()` उपयोग करें या नतीजे storage में लिखें।

**Quiz:** Spark job क्या शुरू करता है?

- [x] count() या write जैसा action
- [ ] DataFrame variable बनाना
- [ ] pyspark import करना
- [ ] Comment टाइप करना

*Answer:* count() या write जैसा action. Transformations सिर्फ़ योजना बनाते हैं; action Spark को उसे सचमुच चलाने पर मजबूर करता है।
