Lesson 11 / 25
Partitioned Output and Partition Pruning
Write data partitioned by a column so queries can skip whole folders.
One folder per value
df.write.partitionBy("order_date").parquet(path) creates one subfolder per value, named order_date=2026-09-01, and so on. A query that filters on that column, such as WHERE order_date = '2026-09-02', reads only the matching folder(s): partition pruning. Choose partition columns with low cardinality that queries filter on, typically a date. Partitioning by a high-cardinality column (like customer_id) creates millions of tiny files, which is worse than not partitioning. Combine with a sensible number of output files per partition.
Writing partitioned Parquet, run
I ran this on Apache Spark 4.0.0 (PySpark, local mode, in the official Docker image). Three dates give three subfolders. Reading with a filter on order_date touches only the matching folder.
import os
path = "/tmp/orders_pq"
orders.write.mode("overwrite").partitionBy("order_date").parquet(path)
print(sorted(d for d in os.listdir(path) if d.startswith("order_date")))
Output:
['order_date=2026-09-01', 'order_date=2026-09-02', 'order_date=2026-09-03']
Look for PartitionFilters in the plan
Run df.explain() and look at the FileScan parquet line: a PartitionFilters entry shows that pruning happened, and PushedFilters shows predicates handed to the Parquet reader.
Quick check: Which is a good partition column for event data?
- the full timestamp to the millisecond
- user_id with millions of values
- a random UUID
- event_date
Answer
event_date — A low-cardinality, frequently filtered column gives useful pruning without a flood of tiny files.