Lesson 21 / 25

Cluster Managers and Platforms

Choose where Spark runs: standalone, YARN, Kubernetes or a managed platform.

Where executors come from

Spark runs on several cluster managers: standalone (Spark's own, simple), YARN (Hadoop clusters), Kubernetes (executors as pods; popular for cloud-native stacks), plus managed platforms that hide the cluster: Databricks, Amazon EMR (and EMR Serverless), Google Dataproc, Azure Synapse/HDInsight. Managed platforms add autoscaling, notebooks, job schedulers and security, trading some control and cost for less operational work. Deploy modes: in client mode the driver runs where you submit (good for notebooks and debugging); in cluster mode it runs inside the cluster (good for production batch jobs). Separate compute from storage by keeping data in object storage (S3, GCS, ADLS) so clusters can be ephemeral.

Run it reliably

Deployment choices, testing, monitoring and table formats turn a notebook job into a dependable pipeline.

Three steps: deploy, test, observe.
Figure 7.1 — Deploy, test and observe.

Choosing a platform

A simple starting guide.

Existing Hadoop/YARN estate            -> Spark on YARN
Cloud-native, containers, one platform  -> Spark on Kubernetes
Want notebooks + autoscale + less ops   -> Databricks / EMR / Dataproc
Small, spiky batch jobs                  -> serverless Spark (e.g. EMR Serverless)
Learning / unit tests                    -> local[*] on a laptop or Docker

Quick check: Why keep data in object storage and not on the cluster?

  • Object storage is always faster than disks
  • So clusters can be temporary and scaled independently of storage
  • Spark cannot read local files
  • It removes the need for backups
Answer

So clusters can be temporary and scaled independently of storage — Separating compute from storage lets you start, resize and stop clusters without moving data.