Lesson 25 / 25

Revision: Cheat Sheet and Self-Check

Review Airflow's concepts, patterns and practices from the whole course.

Cheat sheet

Concepts: DAG, task, operator, DAG run, task instance; scheduler, executor, workers, metadata DB, API server. Writing: with DAG/@dag, @task (TaskFlow), >> and lists, schedule, start_date, catchup; data interval means a run covers a past interval. Data: XCom for small values (pass locations), Jinja ({{ ds }}), Connections/secrets backend. Control: @task.branch, trigger rules (all_done, one_failed, none_failed_min_one_success), sensors (reschedule/deferrable + timeout), dynamic mapping .expand(). Design: idempotent, date-partitioned, ELT, lightweight tasks, safe backfills. Ops: retries + backoff + timeouts + callbacks, pools, monitoring + data-quality checks. Deploy: Git, CI with DagBag import test, pinned versions, PostgreSQL, Celery/Kubernetes executors, managed services.

Questions interviewers ask

Be ready to explain: what a DAG is and how the scheduler works, why tasks must be idempotent, what a data interval and catchup mean, why XComs should stay small, the difference between poke and deferrable sensors, how trigger rules work after a branch, and how you would deploy and test DAG changes safely.

Quick check: A daily DAG with logical date 5 October covers which period?

  • Only the minute it starts
  • 6 Oct to 7 Oct
  • 5 Oct 00:00 to 6 Oct 00:00, running after it ends
  • The whole month
Answer

5 Oct 00:00 to 6 Oct 00:00, running after it ends — The data interval is the day starting at the logical date; the run starts once the interval is over.

Quick check: Everything after a branch is skipped, including the join task. What is the likely fix?

  • Set the join task's trigger_rule to none_failed_min_one_success
  • Delete the branch
  • Increase retries
  • Use datetime.now()
Answer

Set the join task's trigger_rule to none_failed_min_one_success — The default all_success rule skips a join when one upstream path was skipped.

Quick check: Which approach to passing a 5 GB DataFrame between tasks is correct?

  • Email it
  • Return the DataFrame so it is stored in the metadata DB
  • Print it in the logs
  • Write it to object storage and pass the path through XCom
Answer

Write it to object storage and pass the path through XCom — XComs are for small values; large data belongs in external storage.