# explain() से Plans पढ़ना और Adaptive Query Execution — Apache Spark: DataFrames से Big Data Processing

Source: https://www.geekswithgeeks.com/hi/spark/ex-explain-aqe

> देखने के लिए explain() उपयोग करें कि Spark query कैसे चलाएगा और जानें कि AQE run के समय क्या करता है।

## ट्यून करने से पहले plan देखें

`df.explain()` **physical plan** छापता है: Spark query को असल में कैसे चलाएगा। इसे नीचे (डेटा स्रोत) से ऊपर (अंतिम नतीजा) की ओर पढ़ें। `Exchange` (shuffles), join प्रकार (`BroadcastHashJoin`, `SortMergeJoin`), scan विवरण (`PushedFilters`, `PartitionFilters`) और `AdaptiveSparkPlan` देखें। हाल के Spark versions में **Adaptive Query Execution (AQE)** डिफ़ॉल्ट रूप से चालू है: वह असली आँकड़ों से run के समय plan को फिर अनुकूलित करता है, जैसे **छोटे shuffle partitions को मिलाना**, एक ओर छोटा निकलने पर **sort-merge join को broadcast join में बदलना**, और **skewed partitions को बाँटना**। `explain("formatted")` ज़्यादा पठनीय रूप देता है, और Spark UI वही plan metrics के साथ दिखाता है।

## Broadcast join plan, चलाकर

मैंने यह Apache Spark 4.0.0 (PySpark, local mode, official Docker image में) पर चलाया। छोटे `cust` DataFrame को `F.broadcast()` में लपेटने से Spark उसे हर executor तक भेजता है (`BroadcastExchange`) और `BroadcastHashJoin` से join करता है, जिससे बड़ी ओर का shuffle बचता है। `AdaptiveSparkPlan` दिखाता है कि AQE सक्रिय है; `isFinalPlan=false` है क्योंकि अब तक कुछ चला नहीं।

```python
orders.join(F.broadcast(cust), "customer").explain()
```

Output:

```
== Physical Plan ==
AdaptiveSparkPlan isFinalPlan=false
+- Project [customer#1, order_id#0L, city#2, amount#3, order_date#5, tier#72]
   +- BroadcastHashJoin [customer#1], [customer#71], Inner, BuildRight, false
      :- Project [order_id#0L, customer#1, city#2, amount#3, cast(order_date#4 as date) AS order_date#5]
      :  +- Filter isnotnull(customer#1)
      :     +- Scan ExistingRDD[order_id#0L,customer#1,city#2,amount#3,order_date#4]
      +- BroadcastExchange HashedRelationBroadcastMode(List(input[0, string, false]),false), [plan_id=587]
         +- Filter isnotnull(customer#71)
            +- Scan ExistingRDD[customer#71,tier#72]
```

## विशाल tables पर SortMergeJoin पर नज़र रखें

`SortMergeJoin` दोनों ओर shuffle और sort करता है। एक ओर पर्याप्त छोटी हो तो broadcast join बहुत सस्ता है। AQE अपने आप बदल सकता है, पर hints (`F.broadcast`) आपका इरादा स्पष्ट करते हैं।

**Quiz:** कौन-सा plan operator shuffle दर्शाता है?

- [ ] Scan
- [ ] Filter
- [ ] Project
- [x] Exchange

*Answer:* Exchange. Exchange वह जगह है जहाँ डेटा partitions या मशीनों में पुनर्वितरित होता है।
