# What Spark Is — Apache Spark: Big Data Processing with DataFrames

Source: https://www.geekswithgeeks.com/en/spark/sp-what-is

> Explain Spark as a distributed data processing engine and where it fits next to databases and Hadoop.

## Processing too big for one machine

**Apache Spark** is an open-source engine for processing large datasets in parallel across many machines (or many cores on one). You describe what you want with high-level APIs, in Python (**PySpark**), SQL, Scala, Java or R, and Spark splits the data into **partitions**, schedules **tasks**, handles failures and moves data between machines when needed. It keeps intermediate data **in memory** where possible, which made it much faster than the older Hadoop MapReduce for iterative work. Spark covers batch ETL, SQL analytics, machine learning (MLlib) and stream processing in one engine. It is not a database or a storage system: it reads data from files, object storage, databases or message queues, and writes results back.

## Driver plans, executors work

Your program (the driver) builds a plan; a cluster manager starts executors; the executors run tasks on partitions of data in parallel.

![Four parts: driver, cluster manager, executors, data partitions.](assets/figures/spark/section-1-map.svg) — Figure 1.1 — Driver, cluster manager, executors and partitions.

## A team sorting a mountain of letters

One clerk would take weeks to sort a mountain of letters. A manager splits it into bags, hands bags to many clerks, collects partial results and merges them. The manager is the driver, the clerks are executors and the bags are partitions.

## Do not use Spark for small data

If your data fits comfortably in memory on one machine (a few GB), pandas, DuckDB or a database is simpler and often faster. Spark adds startup time and complexity that only pay off at larger scale.

**Quiz:** What is Spark mainly used for?

- [ ] Designing web pages
- [ ] Storing data permanently like a database
- [x] Parallel processing of large datasets across many cores or machines
- [ ] Managing passwords

*Answer:* Parallel processing of large datasets across many cores or machines. Spark is a compute engine that reads from and writes to separate storage.
