Lesson 18 / 25
Keep Airflow Lightweight
Delegate heavy computation to Spark, dbt or the warehouse and keep tasks small and focused.
Orchestrate, do not compute
Run heavy work where it belongs: submit a Spark job, run a dbt build, execute SQL inside the warehouse, or start a Kubernetes pod (KubernetesPodOperator) with its own resources and dependencies, then let Airflow track the outcome. This avoids overloading workers, keeps Python dependencies of different jobs from clashing, and scales independently. Make tasks small and focused (one clear job each) so failures are easy to understand and retries re-do little work, and prefer ELT (load raw data first, transform inside the warehouse) where practical.
A task that only orchestrates
The heavy transformation happens in the warehouse via a dbt build; Airflow just starts it and watches. Illustrative.
dbt_build = BashOperator(
task_id="dbt_build",
bash_command="cd /opt/dbt && dbt build --select tag:daily --vars '{run_date: {{ ds }}}'",
execution_timeout=timedelta(hours=1),
retries=1,
)Quick check: Where should heavy data transformation usually run?
- In systems built for it, such as Spark or the warehouse
- Inside the Airflow scheduler
- Inside XCom
- In the browser
Answer
In systems built for it, such as Spark or the warehouse — Airflow coordinates work; specialised engines process the data.