Lesson 21 / 25

Monitoring, Logs and SLAs

Watch pipeline health with the UI, logs, metrics and alerts, and track lateness.

Know before the business does

The UI shows DAG runs, task states, durations, logs and the graph; learn to read the Grid and Graph views. For production add metrics (Airflow can emit StatsD/OpenTelemetry metrics for scheduler lag, task durations and failures) to a dashboard such as Grafana, centralise logs (remote logging to object storage), and set alerts for failures, long-running tasks and runs that did not start. Track freshness of the data your pipelines produce, because "the DAG succeeded" is not the same as "the data is correct and on time". Add data-quality checks (row counts, nulls, duplicates) as tasks that fail the run when expectations are broken.

Lateness arithmetic, run

I ran this: a report scheduled for 06:00 finished at 06:47 against a 30-minute target, so it was 17 minutes later than allowed.

from datetime import datetime, timedelta
sched = datetime(2026, 10, 2, 6, 0); finished = datetime(2026, 10, 2, 6, 47); target = timedelta(minutes=30)
print("late by", finished - sched - target)

Output:

late by 0:17:00

A data-quality task

Fail the run when the day's row count is implausible. warehouse_count stands for your own query helper. Illustrative.

@task
def check_rows(ds=None) -> None:
    n = warehouse_count(f"SELECT COUNT(*) FROM analytics.daily_orders WHERE order_date = '{ds}'")
    if n < 1000:
        raise ValueError(f"Only {n} rows for {ds}; expected at least 1000")

Quick check: Why check data quality inside the pipeline?

  • Quality checks make DAGs fail more often by design only
  • A successful DAG run can still produce wrong or missing data
  • Airflow cannot run SQL
  • It replaces monitoring
Answer

A successful DAG run can still produce wrong or missing data — Task success only means the code ran; explicit checks confirm the data meets expectations.