Lesson 3 / 25
Running Spark and Creating a SparkSession
Run PySpark locally or in Docker and create a SparkSession.
Local mode is enough to learn
To learn, run Spark in local mode, where the driver and executors are threads in one process: pip install pyspark (requires a Java runtime) or use the official Docker image apache/spark. Every program starts by creating a SparkSession, the entry point for DataFrames and SQL; master("local[2]") means two worker threads. In production you submit applications to a cluster manager with spark-submit, and the master setting comes from the submit command, not from code. Spark's default log output is verbose, so set the log level to ERROR while learning.
Session and a first DataFrame, run
I ran this on Apache Spark 4.0.0 (PySpark, local mode, in the official Docker image). The version line confirms the engine. The six-row orders DataFrame (one amount is null on purpose) is used in the next sections.
from pyspark.sql import SparkSession, functions as F, Window
spark = (SparkSession.builder.master("local[2]").appName("demo")
.config("spark.sql.shuffle.partitions", "4").config("spark.ui.enabled", "false").getOrCreate())
spark.sparkContext.setLogLevel("ERROR")
print("version", spark.version)
orders = spark.createDataFrame([
(1, "asha", "pune", 120.0, "2026-09-01"), (2, "ravi", "delhi", 80.0, "2026-09-01"),
(3, "asha", "pune", 200.0, "2026-09-02"), (4, "meera", "delhi", 50.0, "2026-09-02"),
(5, "ravi", "delhi", 300.0, "2026-09-03"), (6, "kiran", "mumbai", None, "2026-09-03"),
], ["order_id", "customer", "city", "amount", "order_date"])
orders = orders.withColumn("order_date", F.to_date("order_date"))
Output:
version 4.0.0
Submitting to a cluster (illustrative)
In production the same script is submitted with spark-submit. Option values depend on your cluster.
spark-submit \
--master yarn --deploy-mode cluster \
--num-executors 10 --executor-cores 4 --executor-memory 8g \
--conf spark.sql.shuffle.partitions=400 \
daily_orders_job.py --date 2026-10-01Quick check: What is the entry point for DataFrames and SQL in PySpark?
- pandas.read()
- SparkPlug
- DataBase()
- SparkSession
Answer
SparkSession — SparkSession is created first and used to read data and run SQL.