Lesson 5 / 25

Selecting, Filtering and Columns

Use select, filter, withColumn and column expressions.

Columns are expressions

DataFrames are immutable: every operation returns a new DataFrame. Refer to columns with F.col("name") (where F is pyspark.sql.functions) or df["name"]. select chooses or computes columns, filter (alias where) keeps rows matching a condition, withColumn adds or replaces a column, drop removes one, and orderBy sorts. Combine conditions with &, | and ~, each condition in its own parentheses, because Python operator precedence would otherwise mix them up. Prefer built-in functions in pyspark.sql.functions over Python UDFs.

Select, filter, group, join, rank

Most Spark code is a chain of transformations over DataFrames; each step returns a new DataFrame.

Five steps: select, filter, aggregate, join, window.
Figure 2.1 — Select, filter, aggregate, join and window.

Filter and select, run

I ran this on Apache Spark 4.0.0 (PySpark, local mode, in the official Docker image). Only orders with amount above 100 are kept; the null amount is dropped by the comparison.

orders.filter(F.col("amount") > 100).select("order_id", "customer", "amount").orderBy("order_id").show()

Output:

+--------+--------+------+
|order_id|customer|amount|
+--------+--------+------+
|       1|    asha| 120.0|
|       3|    asha| 200.0|
|       5|    ravi| 300.0|
+--------+--------+------+

Parentheses around each condition

Write (F.col("a") > 1) & (F.col("b") == "x"), not F.col("a") > 1 & F.col("b") == "x". The second form fails or gives wrong results because & binds tighter than the comparisons.

Quick check: Why does `df.filter(...)` not modify `df` itself?

  • Filters are ignored
  • DataFrames are immutable; each operation returns a new DataFrame
  • It deletes the data
  • Spark forbids filters
Answer

DataFrames are immutable; each operation returns a new DataFrame — Assign the result to a variable to keep the transformed DataFrame.