# File Formats: CSV, JSON, Parquet — Apache Spark: Big Data Processing with DataFrames

Source: https://www.geekswithgeeks.com/en/spark/sql-file-formats

> Read and write the common formats and prefer columnar formats for analytics.

## Row formats vs columnar

**CSV** and **JSON** are row-oriented text formats: easy to produce and inspect, but large and slow, and CSV carries no types. **Parquet** and **ORC** are **columnar**, compressed binary formats that store the schema, let Spark read only the columns it needs (**column pruning**) and skip chunks using statistics (**predicate pushdown**). For analytics, convert raw CSV/JSON to Parquet early. When reading CSV or JSON, give Spark an explicit **schema** instead of `inferSchema`, which scans the data an extra time and can guess types wrongly. Write modes: `overwrite`, `append`, `ignore`, `errorifexists` (default).

## Reading CSV with an explicit schema (illustrative)

A DDL string is the shortest way to give a schema. `PERMISSIVE` mode keeps bad rows with nulls; `FAILFAST` stops on the first bad row.

```python
schema = "order_id LONG, customer STRING, city STRING, amount DOUBLE, order_date DATE"
df = (spark.read.format("csv")
      .option("header", True)
      .option("mode", "FAILFAST")
      .schema(schema)
      .load("s3a://lake/raw/orders/2026-10-01/*.csv"))

df.write.mode("overwrite").parquet("s3a://lake/clean/orders/")
```

## Avoid one giant CSV

A single large gzip-compressed CSV cannot be split, so one task reads all of it. Store data as many moderately sized Parquet files (roughly 128 MB to 1 GB each).

**Quiz:** Why is Parquet usually better than CSV for analytics?

- [x] It is columnar, compressed and stores the schema, so Spark reads less data
- [ ] It can be opened in a text editor
- [ ] It has no types
- [ ] It is slower to read

*Answer:* It is columnar, compressed and stores the schema, so Spark reads less data. Columnar layout allows column pruning and predicate pushdown.
