# Partitioned Output and Partition Pruning — Apache Spark: Big Data Processing with DataFrames

Source: https://www.geekswithgeeks.com/en/spark/sql-partitioned-output

> Write data partitioned by a column so queries can skip whole folders.

## One folder per value

`df.write.partitionBy("order_date").parquet(path)` creates one subfolder per value, named `order_date=2026-09-01`, and so on. A query that filters on that column, such as `WHERE order_date = '2026-09-02'`, reads only the matching folder(s): **partition pruning**. Choose partition columns with **low cardinality** that queries filter on, typically a date. Partitioning by a high-cardinality column (like `customer_id`) creates millions of tiny files, which is worse than not partitioning. Combine with a sensible number of output files per partition.

## Writing partitioned Parquet, run

I ran this on Apache Spark 4.0.0 (PySpark, local mode, in the official Docker image). Three dates give three subfolders. Reading with a filter on `order_date` touches only the matching folder.

```python
import os
path = "/tmp/orders_pq"
orders.write.mode("overwrite").partitionBy("order_date").parquet(path)
print(sorted(d for d in os.listdir(path) if d.startswith("order_date")))
```

Output:

```
['order_date=2026-09-01', 'order_date=2026-09-02', 'order_date=2026-09-03']
```

## Look for PartitionFilters in the plan

Run `df.explain()` and look at the `FileScan parquet` line: a `PartitionFilters` entry shows that pruning happened, and `PushedFilters` shows predicates handed to the Parquet reader.

**Quiz:** Which is a good partition column for event data?

- [ ] the full timestamp to the millisecond
- [ ] user_id with millions of values
- [ ] a random UUID
- [x] event_date

*Answer:* event_date. A low-cardinality, frequently filtered column gives useful pruning without a flood of tiny files.
