# What Airflow Is and Is Not — Apache Airflow: Orchestrate Data Pipelines

Source: https://www.geekswithgeeks.com/en/airflow/af-what-is

> Describe Airflow as a workflow orchestrator and know what it should not be used for.

## An orchestrator, not a data engine

**Apache Airflow** is an open-source platform to author, schedule and monitor **workflows**. A workflow is a **DAG** (directed acyclic graph) of **tasks** with dependencies, defined in ordinary Python. Airflow decides **when** things run, in **what order**, retries failures and records the history in a UI. It is an **orchestrator**: it should tell Spark, a warehouse, dbt or an API to do the heavy work, not process large datasets inside its own workers. It is built for **batch** workflows with a clear schedule, not for real-time streaming.

## Define in Python, run on a schedule

You write workflows as Python code; Airflow schedules them, runs the tasks in order and shows you what happened.

![Four stages: define, schedule, execute, observe.](assets/figures/airflow/section-1-map.svg) — Figure 1.1 — Define, schedule, execute and observe.

## A railway control room

The control room sets timetables, signals which train may proceed and records delays. It does not pull the trains itself; the engines do. Airflow sets and monitors the timetable; your systems do the pulling.

## Pick the right tool

For event streams use Kafka or Flink; for a single cron job on one server a plain cron entry may be enough. Choose Airflow when you have many dependent steps, need retries, history and visibility.

**Quiz:** What is Airflow best at?

- [ ] Real-time stream processing
- [ ] Processing huge datasets inside its own workers
- [x] Scheduling and monitoring batch workflows of dependent tasks
- [ ] Serving web pages

*Answer:* Scheduling and monitoring batch workflows of dependent tasks. Airflow orchestrates; the heavy lifting should happen in systems built for it.
