# Architecture: Scheduler, Executor, Workers, Database — Apache Airflow: Orchestrate Data Pipelines

Source: https://www.geekswithgeeks.com/en/airflow/af-architecture

> Name the main Airflow components and what each does.

## The moving parts

A **DAG processor** parses your Python files into DAG definitions. The **scheduler** decides which DAG runs and tasks are ready. An **executor** sends those tasks to **workers** that run them (locally, on Celery workers or as Kubernetes pods). The **metadata database** (PostgreSQL or MySQL in production) stores DAG runs, task states and history. A **web server / API server** serves the UI. Airflow 3 introduced a separate API server and a task execution interface so workers talk to the API rather than the database directly. Details vary between versions, so check the documentation for yours.

## Components at a glance

Read it as the path of one scheduled run.

```text
dags/*.py --> DAG processor --parses--> metadata DB (serialized DAGs)
                                 |
Scheduler --creates DAG runs & queues ready tasks--> Executor
                                 |                        |
                                 |             Worker(s) run the task code
                                 v                        |
API server / UI <------ task states, logs <---------------+
```

## Keep DAG parsing fast

The scheduler re-parses DAG files regularly. Avoid slow work (database calls, API requests, heavy imports) at the top level of a DAG file; do it inside tasks.

**Quiz:** Which component decides which tasks are ready to run?

- [x] The scheduler
- [ ] The browser
- [ ] The metadata cache of the OS
- [ ] The Python interpreter alone

*Answer:* The scheduler. The scheduler evaluates schedules and dependencies and queues tasks for the executor.
