# Running Spark and Creating a SparkSession — Apache Spark: Big Data Processing with DataFrames

Source: https://www.geekswithgeeks.com/en/spark/sp-install-session

> Run PySpark locally or in Docker and create a SparkSession.

## Local mode is enough to learn

To learn, run Spark in **local mode**, where the driver and executors are threads in one process: `pip install pyspark` (requires a Java runtime) or use the official Docker image `apache/spark`. Every program starts by creating a **`SparkSession`**, the entry point for DataFrames and SQL; `master("local[2]")` means two worker threads. In production you submit applications to a cluster manager with `spark-submit`, and the master setting comes from the submit command, not from code. Spark's default log output is verbose, so set the log level to `ERROR` while learning.

## Session and a first DataFrame, run

I ran this on Apache Spark 4.0.0 (PySpark, local mode, in the official Docker image). The version line confirms the engine. The six-row `orders` DataFrame (one amount is null on purpose) is used in the next sections.

```python
from pyspark.sql import SparkSession, functions as F, Window
spark = (SparkSession.builder.master("local[2]").appName("demo")
         .config("spark.sql.shuffle.partitions", "4").config("spark.ui.enabled", "false").getOrCreate())
spark.sparkContext.setLogLevel("ERROR")
print("version", spark.version)

orders = spark.createDataFrame([
    (1, "asha", "pune", 120.0, "2026-09-01"), (2, "ravi", "delhi", 80.0, "2026-09-01"),
    (3, "asha", "pune", 200.0, "2026-09-02"), (4, "meera", "delhi", 50.0, "2026-09-02"),
    (5, "ravi", "delhi", 300.0, "2026-09-03"), (6, "kiran", "mumbai", None, "2026-09-03"),
], ["order_id", "customer", "city", "amount", "order_date"])
orders = orders.withColumn("order_date", F.to_date("order_date"))

```

Output:

```
version 4.0.0
```

## Submitting to a cluster (illustrative)

In production the same script is submitted with `spark-submit`. Option values depend on your cluster.

```bash
spark-submit \
  --master yarn --deploy-mode cluster \
  --num-executors 10 --executor-cores 4 --executor-memory 8g \
  --conf spark.sql.shuffle.partitions=400 \
  daily_orders_job.py --date 2026-10-01
```

**Quiz:** What is the entry point for DataFrames and SQL in PySpark?

- [ ] pandas.read()
- [ ] SparkPlug
- [ ] DataBase()
- [x] SparkSession

*Answer:* SparkSession. SparkSession is created first and used to read data and run SQL.
