Lesson 22 / 25

Metrics That Matter

Track steps, tokens, cost, latency and the reason each run ended.

One record per run

Write one structured record at the end of every run: steps used, input and output tokens, cost, wall time, the stop reason, and the tool call counts. Charting these shows what is typical, which runs are outliers and whether a prompt change made things better or worse.

See it, then test it

You cannot tune a loop you cannot see, and you cannot trust one you cannot test.

Three steps: record, replay, assert.
Figure 7.1 — Record, replay and assert.

A run summary record

Log it as one JSON line so tools can filter and chart it. Do not include secrets or full user content.

{"run_id":"r-8841","steps":9,"input_tokens":61200,"output_tokens":2400,
 "cost_usd":0.22,"seconds":41.7,"stop_reason":"end_turn",
 "tools":{"search":4,"read_file":3,"write_file":1}}

Watch the tail, not the average

Averages hide the rare run that costs fifty times more. Track the 95th percentile and the maximum, and alert on runs that hit a hard limit.

Quick check: Why log the stop reason for every run?

  • To see how runs end and which limits are being hit
  • It makes runs faster
  • Logs are required by law
  • It saves tokens
Answer

To see how runs end and which limits are being hit — Stop reasons reveal patterns such as frequent token-budget stops that point to a context problem.