Lesson 24 / 25
Case Study: Budgeting a Research Agent
Design limits and budgets for an agent that searches the web and writes a report.
Reasoning from the task
A research agent searches, reads pages and writes a 1-page report. Searches are cheap but page text is large, so clip every page to about 4,000 tokens. Allow 20 steps, 150,000 total tokens and 3 minutes. Stop early on a repeated search. At 90% of any limit, make a tool-free wrap-up call that reports what was found. Save a checkpoint after each step and log one metrics record per run.
A safe loop end to end
Combine limits, repeat detection, context control, graceful endings and metrics into one design.
The design on one page
Each line maps to a section of this course. Tune the numbers from real runs.
Limits 20 steps, 150k tokens, 180 s, $1.00 (Sec 2)
Guards RepeatGuard, 3 consecutive tool failures (Sec 2, 5)
Context clip pages to ~4k tokens, summarise at 60% (Sec 3, 4)
Caching stable system prompt + tool defs first (Sec 3)
Ending wrap-up call at 90%, handover on failure (Sec 6)
Visibility one JSON record per run, stop reason logged (Sec 7)Quick check: Why clip each fetched page in the research agent?
- Pages are never useful
- Large page text is resent every turn and inflates cost
- The API forbids pages
- Clipping speeds up the network
Answer
Large page text is resent every turn and inflates cost — Big tool results are resent on every later turn, so bounding them controls total cost.