# Transformations, Actions और Lazy Evaluation — Apache Spark: DataFrames से Big Data Processing

Source: https://www.geekswithgeeks.com/hi/spark/ex-lazy

> Transformations और actions में फ़र्क़ करें और समझें कि action तक कुछ क्यों नहीं चलता।

## नतीजा माँगने तक कुछ नहीं होता

**Transformations** (`select`, `filter`, `withColumn`, `join`, `groupBy`) गणना का वर्णन करते हैं और तुरंत लौटते हैं; वे सिर्फ़ logical plan में जुड़ते हैं। **Actions** (`count`, `show`, `collect`, `write`, `first`, `take`) Spark को योजना चलाने और नतीजा लौटाने या सहेजने पर मजबूर करते हैं। यह **lazy evaluation** Spark को आपकी पूरी pipeline देखने और एक के रूप में अनुकूलित करने देता है, जैसे filter को join से पहले ले जाना या सिर्फ़ ज़रूरी columns पढ़ना। क़ीमत यह है कि errors देर से (action पर) दिखते हैं, और action दो बार बुलाने पर cache न हो तो सब कुछ दोबारा गणना होता है।

## पहले योजना, फिर निष्पादन

Transformations योजना बनाते हैं; action उसे अनुकूलित करता है, shuffles पर stages में काटता है और tasks चलाता है।

![चार विचार: lazy, shuffle, योजना, partitions।](assets/figures/spark/section-4-map.svg) — चित्र 4.1 — Lazy, shuffle, योजना और partitions।

## दस लाख rows, एक job, चलाकर

मैंने यह Apache Spark 4.0.0 (PySpark, local mode, official Docker image में) पर चलाया। `filter` और `withColumn` पंक्तियाँ बिना काम किए तुरंत लौटती हैं; सिर्फ़ `count()` job चलाता है। 0 से 999,999 की आधी संख्याएँ सम हैं।

```python
big = spark.range(0, 1000000)
t = big.filter("id % 2 = 0").withColumn("sq", F.col("id") * 2)   # nothing runs yet
print("no job yet; count =", t.count())          # the action triggers the job
```

Output:

```
no job yet; count = 500000
```

## विकास के दौरान take() या limit() उपयोग करें

तेज़ प्रतिक्रिया के लिए तर्क को छोटे नमूने (`df.limit(1000)`) पर परखें, फिर पूरे डेटा पर चलाएँ।

**Quiz:** इनमें से कौन-सा action है?

- [ ] filter()
- [x] count()
- [ ] select()
- [ ] withColumn()

*Answer:* count(). count() निष्पादन कराता है; बाक़ी सिर्फ़ योजना बढ़ाते हैं।
