Lesson 21 / 25
Cluster Managers and Platforms
Choose where Spark runs: standalone, YARN, Kubernetes or a managed platform.
Where executors come from
Spark runs on several cluster managers: standalone (Spark's own, simple), YARN (Hadoop clusters), Kubernetes (executors as pods; popular for cloud-native stacks), plus managed platforms that hide the cluster: Databricks, Amazon EMR (and EMR Serverless), Google Dataproc, Azure Synapse/HDInsight. Managed platforms add autoscaling, notebooks, job schedulers and security, trading some control and cost for less operational work. Deploy modes: in client mode the driver runs where you submit (good for notebooks and debugging); in cluster mode it runs inside the cluster (good for production batch jobs). Separate compute from storage by keeping data in object storage (S3, GCS, ADLS) so clusters can be ephemeral.
Run it reliably
Deployment choices, testing, monitoring and table formats turn a notebook job into a dependable pipeline.
Choosing a platform
A simple starting guide.
Existing Hadoop/YARN estate -> Spark on YARN
Cloud-native, containers, one platform -> Spark on Kubernetes
Want notebooks + autoscale + less ops -> Databricks / EMR / Dataproc
Small, spiky batch jobs -> serverless Spark (e.g. EMR Serverless)
Learning / unit tests -> local[*] on a laptop or DockerQuick check: Why keep data in object storage and not on the cluster?
- Object storage is always faster than disks
- So clusters can be temporary and scaled independently of storage
- Spark cannot read local files
- It removes the need for backups
Answer
So clusters can be temporary and scaled independently of storage — Separating compute from storage lets you start, resize and stop clusters without moving data.