Lesson 1 / 25

What Spark Is

Explain Spark as a distributed data processing engine and where it fits next to databases and Hadoop.

Processing too big for one machine

Apache Spark is an open-source engine for processing large datasets in parallel across many machines (or many cores on one). You describe what you want with high-level APIs, in Python (PySpark), SQL, Scala, Java or R, and Spark splits the data into partitions, schedules tasks, handles failures and moves data between machines when needed. It keeps intermediate data in memory where possible, which made it much faster than the older Hadoop MapReduce for iterative work. Spark covers batch ETL, SQL analytics, machine learning (MLlib) and stream processing in one engine. It is not a database or a storage system: it reads data from files, object storage, databases or message queues, and writes results back.

Driver plans, executors work

Your program (the driver) builds a plan; a cluster manager starts executors; the executors run tasks on partitions of data in parallel.

Four parts: driver, cluster manager, executors, data partitions.
Figure 1.1 — Driver, cluster manager, executors and partitions.

A team sorting a mountain of letters

One clerk would take weeks to sort a mountain of letters. A manager splits it into bags, hands bags to many clerks, collects partial results and merges them. The manager is the driver, the clerks are executors and the bags are partitions.

Do not use Spark for small data

If your data fits comfortably in memory on one machine (a few GB), pandas, DuckDB or a database is simpler and often faster. Spark adds startup time and complexity that only pay off at larger scale.

Quick check: What is Spark mainly used for?

  • Designing web pages
  • Storing data permanently like a database
  • Parallel processing of large datasets across many cores or machines
  • Managing passwords
Answer

Parallel processing of large datasets across many cores or machines — Spark is a compute engine that reads from and writes to separate storage.