Lesson 11 / 29
Long Tasks: Compaction, Notes and Fresh Starts
Keep a long session coherent when the context fills up.
The context fills, so decide what to keep
Every step adds file contents and command output to the context, so a long task eventually hits the limit and quality drops as old, irrelevant material crowds out the important. Techniques: compaction (the harness summarises the older conversation into a short note and drops the bulk, keeping the goal, decisions made, files touched and what remains); notes or plan files (the agent writes PLAN.md or a to-do list in the repository and re-reads it, so progress survives context resets and a human can see it); clearing noisy tool output once it has been used; and fresh starts (finish one piece, commit, and begin a new session on the next piece with a short hand-off note). Breaking a large goal into independent, verifiable steps is the most reliable remedy: each step has a small context and a clear test.
A plan file the agent keeps updated
A small written plan lets the work survive a context reset and lets a human follow progress.
# PLAN.md (goal: add pagination to GET /orders)
- [x] 1. read shop/orders.py and tests/test_orders.py; note current response shape
- [x] 2. write failing tests: page size, next cursor, last page
- [ ] 3. implement `paginate()` in shop/orders.py
- [ ] 4. update API docs section "Orders"
- [ ] 5. run all tests + lint; summarise the diff
Decisions: cursor = last id (not offset); default page size 20, max 100.
Open question for human: should page size be configurable per client?Quick check: Which is the most reliable way to handle a very large task?
- Use the largest possible prompt
- Give it all at once and hope
- Remove the tests
- Break it into small, independently verifiable steps
Answer
Break it into small, independently verifiable steps — Small steps keep context small and checks clear.