Lesson 11 / 25

Partitioned Output and Partition Pruning

Write data partitioned by a column so queries can skip whole folders.

One folder per value

df.write.partitionBy("order_date").parquet(path) creates one subfolder per value, named order_date=2026-09-01, and so on. A query that filters on that column, such as WHERE order_date = '2026-09-02', reads only the matching folder(s): partition pruning. Choose partition columns with low cardinality that queries filter on, typically a date. Partitioning by a high-cardinality column (like customer_id) creates millions of tiny files, which is worse than not partitioning. Combine with a sensible number of output files per partition.

Writing partitioned Parquet, run

I ran this on Apache Spark 4.0.0 (PySpark, local mode, in the official Docker image). Three dates give three subfolders. Reading with a filter on order_date touches only the matching folder.

import os
path = "/tmp/orders_pq"
orders.write.mode("overwrite").partitionBy("order_date").parquet(path)
print(sorted(d for d in os.listdir(path) if d.startswith("order_date")))

Output:

['order_date=2026-09-01', 'order_date=2026-09-02', 'order_date=2026-09-03']

Look for PartitionFilters in the plan

Run df.explain() and look at the FileScan parquet line: a PartitionFilters entry shows that pruning happened, and PushedFilters shows predicates handed to the Parquet reader.

Quick check: Which is a good partition column for event data?

  • the full timestamp to the millisecond
  • user_id with millions of values
  • a random UUID
  • event_date
Answer

event_date — A low-cardinality, frequently filtered column gives useful pruning without a flood of tiny files.