# Window Functions — Apache Spark: DataFrames से Big Data Processing

Source: https://www.geekswithgeeks.com/hi/spark/df-windows

> Data को सिकोड़े बिना समूहों के भीतर rows को rank करें और running totals निकालें।

## Rows खोए बिना aggregate

**Window function** हर row के लिए संबंधित rows की "window" से मान निकालता है, जो `Window.partitionBy(...)` (समूह), `orderBy(...)` (उसके भीतर का क्रम) और वैकल्पिक frame से परिभाषित होती है। `groupBy` के विपरीत वह हर row रखता है। आम उपयोग: "प्रति समूह शीर्ष N" के लिए `row_number()` / `rank()` / `dense_rank()`, पिछली या अगली row से तुलना के लिए `lag()` / `lead()`, और **running totals** के लिए `sum().over(...)`। `partitionBy` वाली windows डेटा को partition key से shuffle करती हैं, और बिना `partitionBy` वाली window सब कुछ एक partition में खींच लेती है, जो scale नहीं होता।

## प्रति ग्राहक शीर्ष order, चलाकर

मैंने यह Apache Spark 4.0.0 (PySpark, local mode, official Docker image में) पर चलाया। `desc_nulls_last` null amount को आख़िर में रखता है, इसलिए Kiran का अकेला order अब भी null amount के साथ rank 1 है।

```python
w = Window.partitionBy("customer").orderBy(F.col("amount").desc_nulls_last())
orders.withColumn("rank", F.row_number().over(w)).filter("rank = 1").select("customer", "amount").orderBy("customer").show()
```

Output:

```
+--------+------+
|customer|amount|
+--------+------+
|    asha| 200.0|
|   kiran|  NULL|
|   meera|  50.0|
|    ravi| 300.0|
+--------+------+
```

## Running total, चलाकर

मैंने यह Apache Spark 4.0.0 (PySpark, local mode, official Docker image में) पर चलाया। हर row अब तक की संचयी राशि तारीख़ व order क्रम में दिखाती है; null पहले 0.0 से भरा गया, इसलिए आख़िरी दो rows 750.0 पर रहती हैं।

```python
w2 = Window.orderBy("order_date", "order_id")
orders.na.fill({"amount":0.0}).withColumn("running", F.sum("amount").over(w2)).select("order_id", "running").orderBy("order_id").show()
```

Output:

```
+--------+-------+
|order_id|running|
+--------+-------+
|       1|  120.0|
|       2|  200.0|
|       3|  400.0|
|       4|  450.0|
|       5|  750.0|
|       6|  750.0|
+--------+-------+
```

**Quiz:** groupBy और window function में मुख्य अंतर क्या है?

- [ ] कोई अंतर नहीं
- [ ] groupBy हर row रखता है, windows उन्हें सिकोड़ती हैं
- [x] Window function हर input row रखता है
- [ ] Windows सिर्फ़ strings पर चलती हैं

*Answer:* Window function हर input row रखता है. groupBy हर समूह को एक row में सिकोड़ता है; window functions सभी rows में गणना किया column जोड़ती हैं।
