Lesson 2 / 25

Architecture: Scheduler, Executor, Workers, Database

Name the main Airflow components and what each does.

The moving parts

A DAG processor parses your Python files into DAG definitions. The scheduler decides which DAG runs and tasks are ready. An executor sends those tasks to workers that run them (locally, on Celery workers or as Kubernetes pods). The metadata database (PostgreSQL or MySQL in production) stores DAG runs, task states and history. A web server / API server serves the UI. Airflow 3 introduced a separate API server and a task execution interface so workers talk to the API rather than the database directly. Details vary between versions, so check the documentation for yours.

Components at a glance

Read it as the path of one scheduled run.

dags/*.py --> DAG processor --parses--> metadata DB (serialized DAGs)
                                 |
Scheduler --creates DAG runs & queues ready tasks--> Executor
                                 |                        |
                                 |             Worker(s) run the task code
                                 v                        |
API server / UI <------ task states, logs <---------------+

Keep DAG parsing fast

The scheduler re-parses DAG files regularly. Avoid slow work (database calls, API requests, heavy imports) at the top level of a DAG file; do it inside tasks.

Quick check: Which component decides which tasks are ready to run?

  • The scheduler
  • The browser
  • The metadata cache of the OS
  • The Python interpreter alone
Answer

The scheduler — The scheduler evaluates schedules and dependencies and queues tasks for the executor.