# Data Layout, Small Files and Configuration — Apache Spark: Big Data Processing with DataFrames

Source: https://www.geekswithgeeks.com/en/spark/perf-layout-config

> Fix small-file problems, choose file sizes and set the few configs that matter.

## Where most real slowdowns hide

**Small files** (thousands of files of a few KB) make reading slow because every file costs an open call and a task. Avoid them by controlling output partitions before writing (`coalesce`/`repartition`), compacting periodically, or using table formats such as **Delta Lake** or **Apache Iceberg** that compact and manage files. Use **columnar formats** and select only needed columns. A few settings matter most: `spark.sql.shuffle.partitions` (shuffle width; AQE helps), executor `memory` and `cores` (a common rule is 4-5 cores per executor, avoiding giant heaps), `spark.sql.adaptive.enabled` (keep on), and serialization (Kryo for RDD-heavy jobs). Avoid Python UDFs where a built-in function exists, since each row crosses the JVM-Python boundary; if you need custom logic, prefer **pandas UDFs** (vectorised) or native SQL expressions.

## UDF vs built-in, run

I ran this on Apache Spark 4.0.0 (PySpark, local mode, in the official Docker image). Both approaches give the same answer, `ASHA!`, but the built-in `upper`/`concat` version runs inside the JVM and is optimisable, while the Python UDF is opaque to Catalyst and slower at scale.

```python
from pyspark.sql.types import StringType
shout = F.udf(lambda s: s.upper() + "!", StringType())
print(orders.select(shout("customer").alias("u")).orderBy("u").collect()[0][0], orders.select(F.concat(F.upper("customer"), F.lit("!")).alias("b")).orderBy("b").collect()[0][0])
```

Output:

```
ASHA! ASHA!
```

## A few useful configs (illustrative)

Start from defaults, change one thing at a time and measure in the Spark UI. Values depend on cluster size and data.

```python
spark = (SparkSession.builder
    .config("spark.sql.shuffle.partitions", "400")          # AQE can coalesce this
    .config("spark.sql.adaptive.enabled", "true")
    .config("spark.sql.autoBroadcastJoinThreshold", str(50 * 1024 * 1024))
    .config("spark.executor.memory", "8g")
    .config("spark.executor.cores", "4")
    .getOrCreate())
```

**Quiz:** Why avoid thousands of tiny files?

- [x] Each file adds open-and-task overhead, which slows reads
- [ ] Tiny files are illegal
- [ ] They use too little disk
- [ ] Spark cannot read them at all

*Answer:* Each file adds open-and-task overhead, which slows reads. Per-file overhead dominates when files are small, so compact them into larger ones.
