# Cluster Managers and Platforms — Apache Spark: Big Data Processing with DataFrames

Source: https://www.geekswithgeeks.com/en/spark/prod-deploy

> Choose where Spark runs: standalone, YARN, Kubernetes or a managed platform.

## Where executors come from

Spark runs on several cluster managers: **standalone** (Spark's own, simple), **YARN** (Hadoop clusters), **Kubernetes** (executors as pods; popular for cloud-native stacks), plus **managed platforms** that hide the cluster: Databricks, Amazon EMR (and EMR Serverless), Google Dataproc, Azure Synapse/HDInsight. Managed platforms add autoscaling, notebooks, job schedulers and security, trading some control and cost for less operational work. **Deploy modes**: in `client` mode the driver runs where you submit (good for notebooks and debugging); in `cluster` mode it runs inside the cluster (good for production batch jobs). Separate compute from storage by keeping data in object storage (S3, GCS, ADLS) so clusters can be ephemeral.

## Run it reliably

Deployment choices, testing, monitoring and table formats turn a notebook job into a dependable pipeline.

![Three steps: deploy, test, observe.](assets/figures/spark/section-7-map.svg) — Figure 7.1 — Deploy, test and observe.

## Choosing a platform

A simple starting guide.

```text
Existing Hadoop/YARN estate            -> Spark on YARN
Cloud-native, containers, one platform  -> Spark on Kubernetes
Want notebooks + autoscale + less ops   -> Databricks / EMR / Dataproc
Small, spiky batch jobs                  -> serverless Spark (e.g. EMR Serverless)
Learning / unit tests                    -> local[*] on a laptop or Docker
```

**Quiz:** Why keep data in object storage and not on the cluster?

- [ ] Object storage is always faster than disks
- [x] So clusters can be temporary and scaled independently of storage
- [ ] Spark cannot read local files
- [ ] It removes the need for backups

*Answer:* So clusters can be temporary and scaled independently of storage. Separating compute from storage lets you start, resize and stop clusters without moving data.
