Lesson 5 / 25
Selecting, Filtering and Columns
Use select, filter, withColumn and column expressions.
Columns are expressions
DataFrames are immutable: every operation returns a new DataFrame. Refer to columns with F.col("name") (where F is pyspark.sql.functions) or df["name"]. select chooses or computes columns, filter (alias where) keeps rows matching a condition, withColumn adds or replaces a column, drop removes one, and orderBy sorts. Combine conditions with &, | and ~, each condition in its own parentheses, because Python operator precedence would otherwise mix them up. Prefer built-in functions in pyspark.sql.functions over Python UDFs.
Select, filter, group, join, rank
Most Spark code is a chain of transformations over DataFrames; each step returns a new DataFrame.
Filter and select, run
I ran this on Apache Spark 4.0.0 (PySpark, local mode, in the official Docker image). Only orders with amount above 100 are kept; the null amount is dropped by the comparison.
orders.filter(F.col("amount") > 100).select("order_id", "customer", "amount").orderBy("order_id").show()
Output:
+--------+--------+------+ |order_id|customer|amount| +--------+--------+------+ | 1| asha| 120.0| | 3| asha| 200.0| | 5| ravi| 300.0| +--------+--------+------+
Parentheses around each condition
Write (F.col("a") > 1) & (F.col("b") == "x"), not F.col("a") > 1 & F.col("b") == "x". The second form fails or gives wrong results because & binds tighter than the comparisons.
Quick check: Why does `df.filter(...)` not modify `df` itself?
- Filters are ignored
- DataFrames are immutable; each operation returns a new DataFrame
- It deletes the data
- Spark forbids filters
Answer
DataFrames are immutable; each operation returns a new DataFrame — Assign the result to a variable to keep the transformed DataFrame.