# Selecting, Filtering और Columns — Apache Spark: DataFrames से Big Data Processing

Source: https://www.geekswithgeeks.com/hi/spark/df-select-filter

> select, filter, withColumn और column expressions उपयोग करें।

## Columns expressions हैं

DataFrames **immutable** हैं: हर ऑपरेशन नया DataFrame लौटाता है। Columns को `F.col("name")` (जहाँ `F` `pyspark.sql.functions` है) या `df["name"]` से पुकारें। `select` columns चुनता या गणना करता है, `filter` (उपनाम `where`) शर्त पूरी करने वाली rows रखता है, `withColumn` column जोड़ता या बदलता है, `drop` हटाता है, और `orderBy` क्रम देता है। शर्तों को `&`, `|` और `~` से जोड़ें, हर शर्त अपने कोष्ठक में, क्योंकि वरना Python operator precedence उन्हें उलझा देगा। Python UDFs की जगह `pyspark.sql.functions` के built-in functions चुनें।

## Select, filter, group, join, rank

ज़्यादातर Spark कोड DataFrames पर transformations की श्रृंखला है; हर step नया DataFrame लौटाता है।

![पाँच चरण: select, filter, aggregate, join, window।](assets/figures/spark/section-2-map.svg) — चित्र 2.1 — Select, filter, aggregate, join और window।

## Filter और select, चलाकर

मैंने यह Apache Spark 4.0.0 (PySpark, local mode, official Docker image में) पर चलाया। सिर्फ़ 100 से ऊपर के amount वाले orders रखे गए; null amount तुलना से बाहर हो जाता है।

```python
orders.filter(F.col("amount") > 100).select("order_id", "customer", "amount").orderBy("order_id").show()
```

Output:

```
+--------+--------+------+
|order_id|customer|amount|
+--------+--------+------+
|       1|    asha| 120.0|
|       3|    asha| 200.0|
|       5|    ravi| 300.0|
+--------+--------+------+
```

## हर शर्त के चारों ओर कोष्ठक

`(F.col("a") > 1) & (F.col("b") == "x")` लिखें, `F.col("a") > 1 & F.col("b") == "x"` नहीं। दूसरा रूप विफल होता है या ग़लत नतीजे देता है क्योंकि `&`, तुलनाओं से ज़्यादा कसकर बँधता है।

**Quiz:** `df.filter(...)` स्वयं `df` को क्यों नहीं बदलता?

- [ ] Filters अनदेखे होते हैं
- [x] DataFrames immutable हैं; हर ऑपरेशन नया DataFrame लौटाता है
- [ ] यह डेटा हटाता है
- [ ] Spark filters मना करता है

*Answer:* DataFrames immutable हैं; हर ऑपरेशन नया DataFrame लौटाता है. रूपांतरित DataFrame रखने के लिए नतीजे को variable में रखें।
