# Selecting, Filtering and Columns — Apache Spark: Big Data Processing with DataFrames

Source: https://www.geekswithgeeks.com/en/spark/df-select-filter

> Use select, filter, withColumn and column expressions.

## Columns are expressions

DataFrames are **immutable**: every operation returns a new DataFrame. Refer to columns with `F.col("name")` (where `F` is `pyspark.sql.functions`) or `df["name"]`. `select` chooses or computes columns, `filter` (alias `where`) keeps rows matching a condition, `withColumn` adds or replaces a column, `drop` removes one, and `orderBy` sorts. Combine conditions with `&`, `|` and `~`, each condition in its own parentheses, because Python operator precedence would otherwise mix them up. Prefer built-in functions in `pyspark.sql.functions` over Python UDFs.

## Select, filter, group, join, rank

Most Spark code is a chain of transformations over DataFrames; each step returns a new DataFrame.

![Five steps: select, filter, aggregate, join, window.](assets/figures/spark/section-2-map.svg) — Figure 2.1 — Select, filter, aggregate, join and window.

## Filter and select, run

I ran this on Apache Spark 4.0.0 (PySpark, local mode, in the official Docker image). Only orders with amount above 100 are kept; the null amount is dropped by the comparison.

```python
orders.filter(F.col("amount") > 100).select("order_id", "customer", "amount").orderBy("order_id").show()
```

Output:

```
+--------+--------+------+
|order_id|customer|amount|
+--------+--------+------+
|       1|    asha| 120.0|
|       3|    asha| 200.0|
|       5|    ravi| 300.0|
+--------+--------+------+
```

## Parentheses around each condition

Write `(F.col("a") > 1) & (F.col("b") == "x")`, not `F.col("a") > 1 & F.col("b") == "x"`. The second form fails or gives wrong results because `&` binds tighter than the comparisons.

**Quiz:** Why does `df.filter(...)` not modify `df` itself?

- [ ] Filters are ignored
- [x] DataFrames are immutable; each operation returns a new DataFrame
- [ ] It deletes the data
- [ ] Spark forbids filters

*Answer:* DataFrames are immutable; each operation returns a new DataFrame. Assign the result to a variable to keep the transformed DataFrame.
