Lesson 24 / 25

Case Study: Budgeting a Research Agent

Design limits and budgets for an agent that searches the web and writes a report.

Reasoning from the task

A research agent searches, reads pages and writes a 1-page report. Searches are cheap but page text is large, so clip every page to about 4,000 tokens. Allow 20 steps, 150,000 total tokens and 3 minutes. Stop early on a repeated search. At 90% of any limit, make a tool-free wrap-up call that reports what was found. Save a checkpoint after each step and log one metrics record per run.

A safe loop end to end

Combine limits, repeat detection, context control, graceful endings and metrics into one design.

Four layers around the loop core.
Figure 8.1 — Limits, guards, context control and reporting around the loop.

The design on one page

Each line maps to a section of this course. Tune the numbers from real runs.

Limits       20 steps, 150k tokens, 180 s, $1.00        (Sec 2)
Guards       RepeatGuard, 3 consecutive tool failures     (Sec 2, 5)
Context      clip pages to ~4k tokens, summarise at 60%   (Sec 3, 4)
Caching      stable system prompt + tool defs first       (Sec 3)
Ending       wrap-up call at 90%, handover on failure     (Sec 6)
Visibility   one JSON record per run, stop reason logged  (Sec 7)

Quick check: Why clip each fetched page in the research agent?

  • Pages are never useful
  • Large page text is resent every turn and inflates cost
  • The API forbids pages
  • Clipping speeds up the network
Answer

Large page text is resent every turn and inflates cost — Big tool results are resent on every later turn, so bounding them controls total cost.