Lesson 18 / 25

Data Layout, Small Files and Configuration

Fix small-file problems, choose file sizes and set the few configs that matter.

Where most real slowdowns hide

Small files (thousands of files of a few KB) make reading slow because every file costs an open call and a task. Avoid them by controlling output partitions before writing (coalesce/repartition), compacting periodically, or using table formats such as Delta Lake or Apache Iceberg that compact and manage files. Use columnar formats and select only needed columns. A few settings matter most: spark.sql.shuffle.partitions (shuffle width; AQE helps), executor memory and cores (a common rule is 4-5 cores per executor, avoiding giant heaps), spark.sql.adaptive.enabled (keep on), and serialization (Kryo for RDD-heavy jobs). Avoid Python UDFs where a built-in function exists, since each row crosses the JVM-Python boundary; if you need custom logic, prefer pandas UDFs (vectorised) or native SQL expressions.

UDF vs built-in, run

I ran this on Apache Spark 4.0.0 (PySpark, local mode, in the official Docker image). Both approaches give the same answer, ASHA!, but the built-in upper/concat version runs inside the JVM and is optimisable, while the Python UDF is opaque to Catalyst and slower at scale.

from pyspark.sql.types import StringType
shout = F.udf(lambda s: s.upper() + "!", StringType())
print(orders.select(shout("customer").alias("u")).orderBy("u").collect()[0][0], orders.select(F.concat(F.upper("customer"), F.lit("!")).alias("b")).orderBy("b").collect()[0][0])

Output:

ASHA! ASHA!

A few useful configs (illustrative)

Start from defaults, change one thing at a time and measure in the Spark UI. Values depend on cluster size and data.

spark = (SparkSession.builder
    .config("spark.sql.shuffle.partitions", "400")          # AQE can coalesce this
    .config("spark.sql.adaptive.enabled", "true")
    .config("spark.sql.autoBroadcastJoinThreshold", str(50 * 1024 * 1024))
    .config("spark.executor.memory", "8g")
    .config("spark.executor.cores", "4")
    .getOrCreate())

Quick check: Why avoid thousands of tiny files?

  • Each file adds open-and-task overhead, which slows reads
  • Tiny files are illegal
  • They use too little disk
  • Spark cannot read them at all
Answer

Each file adds open-and-task overhead, which slows reads — Per-file overhead dominates when files are small, so compact them into larger ones.