# Revision: Cheat Sheet and Self-Check — Apache Airflow: Orchestrate Data Pipelines

Source: https://www.geekswithgeeks.com/en/airflow/wrap-revision

> Review Airflow's concepts, patterns and practices from the whole course.

## Cheat sheet

**Concepts**: DAG, task, operator, DAG run, task instance; scheduler, executor, workers, metadata DB, API server. **Writing**: `with DAG`/`@dag`, `@task` (TaskFlow), `>>` and lists, `schedule`, `start_date`, `catchup`; data interval means a run covers a past interval. **Data**: XCom for small values (pass locations), Jinja (`{{ ds }}`), Connections/secrets backend. **Control**: `@task.branch`, trigger rules (`all_done`, `one_failed`, `none_failed_min_one_success`), sensors (reschedule/deferrable + timeout), dynamic mapping `.expand()`. **Design**: idempotent, date-partitioned, ELT, lightweight tasks, safe backfills. **Ops**: retries + backoff + timeouts + callbacks, pools, monitoring + data-quality checks. **Deploy**: Git, CI with DagBag import test, pinned versions, PostgreSQL, Celery/Kubernetes executors, managed services.

## Questions interviewers ask

Be ready to explain: what a DAG is and how the scheduler works, why tasks must be idempotent, what a data interval and catchup mean, why XComs should stay small, the difference between poke and deferrable sensors, how trigger rules work after a branch, and how you would deploy and test DAG changes safely.

**Quiz:** A daily DAG with logical date 5 October covers which period?

- [ ] Only the minute it starts
- [ ] 6 Oct to 7 Oct
- [x] 5 Oct 00:00 to 6 Oct 00:00, running after it ends
- [ ] The whole month

*Answer:* 5 Oct 00:00 to 6 Oct 00:00, running after it ends. The data interval is the day starting at the logical date; the run starts once the interval is over.

**Quiz:** Everything after a branch is skipped, including the join task. What is the likely fix?

- [x] Set the join task's trigger_rule to none_failed_min_one_success
- [ ] Delete the branch
- [ ] Increase retries
- [ ] Use datetime.now()

*Answer:* Set the join task's trigger_rule to none_failed_min_one_success. The default all_success rule skips a join when one upstream path was skipped.

**Quiz:** Which approach to passing a 5 GB DataFrame between tasks is correct?

- [ ] Email it
- [ ] Return the DataFrame so it is stored in the metadata DB
- [ ] Print it in the logs
- [x] Write it to object storage and pass the path through XCom

*Answer:* Write it to object storage and pass the path through XCom. XComs are for small values; large data belongs in external storage.
