Lesson 10 / 25

File Formats: CSV, JSON, Parquet

Read and write the common formats and prefer columnar formats for analytics.

Row formats vs columnar

CSV and JSON are row-oriented text formats: easy to produce and inspect, but large and slow, and CSV carries no types. Parquet and ORC are columnar, compressed binary formats that store the schema, let Spark read only the columns it needs (column pruning) and skip chunks using statistics (predicate pushdown). For analytics, convert raw CSV/JSON to Parquet early. When reading CSV or JSON, give Spark an explicit schema instead of inferSchema, which scans the data an extra time and can guess types wrongly. Write modes: overwrite, append, ignore, errorifexists (default).

Reading CSV with an explicit schema (illustrative)

A DDL string is the shortest way to give a schema. PERMISSIVE mode keeps bad rows with nulls; FAILFAST stops on the first bad row.

schema = "order_id LONG, customer STRING, city STRING, amount DOUBLE, order_date DATE"
df = (spark.read.format("csv")
      .option("header", True)
      .option("mode", "FAILFAST")
      .schema(schema)
      .load("s3a://lake/raw/orders/2026-10-01/*.csv"))

df.write.mode("overwrite").parquet("s3a://lake/clean/orders/")

Avoid one giant CSV

A single large gzip-compressed CSV cannot be split, so one task reads all of it. Store data as many moderately sized Parquet files (roughly 128 MB to 1 GB each).

Quick check: Why is Parquet usually better than CSV for analytics?

  • It is columnar, compressed and stores the schema, so Spark reads less data
  • It can be opened in a text editor
  • It has no types
  • It is slower to read
Answer

It is columnar, compressed and stores the schema, so Spark reads less data — Columnar layout allows column pruning and predicate pushdown.