Lesson 11 / 29

Long Tasks: Compaction, Notes and Fresh Starts

Keep a long session coherent when the context fills up.

The context fills, so decide what to keep

Every step adds file contents and command output to the context, so a long task eventually hits the limit and quality drops as old, irrelevant material crowds out the important. Techniques: compaction (the harness summarises the older conversation into a short note and drops the bulk, keeping the goal, decisions made, files touched and what remains); notes or plan files (the agent writes PLAN.md or a to-do list in the repository and re-reads it, so progress survives context resets and a human can see it); clearing noisy tool output once it has been used; and fresh starts (finish one piece, commit, and begin a new session on the next piece with a short hand-off note). Breaking a large goal into independent, verifiable steps is the most reliable remedy: each step has a small context and a clear test.

A plan file the agent keeps updated

A small written plan lets the work survive a context reset and lets a human follow progress.

# PLAN.md  (goal: add pagination to GET /orders)
- [x] 1. read shop/orders.py and tests/test_orders.py; note current response shape
- [x] 2. write failing tests: page size, next cursor, last page
- [ ] 3. implement `paginate()` in shop/orders.py
- [ ] 4. update API docs section "Orders"
- [ ] 5. run all tests + lint; summarise the diff
Decisions: cursor = last id (not offset); default page size 20, max 100.
Open question for human: should page size be configurable per client?

Quick check: Which is the most reliable way to handle a very large task?

  • Use the largest possible prompt
  • Give it all at once and hope
  • Remove the tests
  • Break it into small, independently verifiable steps
Answer

Break it into small, independently verifiable steps — Small steps keep context small and checks clear.