Lesson 18 / 25
Data Layout, Small Files and Configuration
Fix small-file problems, choose file sizes and set the few configs that matter.
Where most real slowdowns hide
Small files (thousands of files of a few KB) make reading slow because every file costs an open call and a task. Avoid them by controlling output partitions before writing (coalesce/repartition), compacting periodically, or using table formats such as Delta Lake or Apache Iceberg that compact and manage files. Use columnar formats and select only needed columns. A few settings matter most: spark.sql.shuffle.partitions (shuffle width; AQE helps), executor memory and cores (a common rule is 4-5 cores per executor, avoiding giant heaps), spark.sql.adaptive.enabled (keep on), and serialization (Kryo for RDD-heavy jobs). Avoid Python UDFs where a built-in function exists, since each row crosses the JVM-Python boundary; if you need custom logic, prefer pandas UDFs (vectorised) or native SQL expressions.
UDF vs built-in, run
I ran this on Apache Spark 4.0.0 (PySpark, local mode, in the official Docker image). Both approaches give the same answer, ASHA!, but the built-in upper/concat version runs inside the JVM and is optimisable, while the Python UDF is opaque to Catalyst and slower at scale.
from pyspark.sql.types import StringType
shout = F.udf(lambda s: s.upper() + "!", StringType())
print(orders.select(shout("customer").alias("u")).orderBy("u").collect()[0][0], orders.select(F.concat(F.upper("customer"), F.lit("!")).alias("b")).orderBy("b").collect()[0][0])
Output:
ASHA! ASHA!
A few useful configs (illustrative)
Start from defaults, change one thing at a time and measure in the Spark UI. Values depend on cluster size and data.
spark = (SparkSession.builder
.config("spark.sql.shuffle.partitions", "400") # AQE can coalesce this
.config("spark.sql.adaptive.enabled", "true")
.config("spark.sql.autoBroadcastJoinThreshold", str(50 * 1024 * 1024))
.config("spark.executor.memory", "8g")
.config("spark.executor.cores", "4")
.getOrCreate())Quick check: Why avoid thousands of tiny files?
- Each file adds open-and-task overhead, which slows reads
- Tiny files are illegal
- They use too little disk
- Spark cannot read them at all
Answer
Each file adds open-and-task overhead, which slows reads — Per-file overhead dominates when files are small, so compact them into larger ones.