Lesson 23 / 29
What a Run Costs: Context Growth and Caching
Estimate cost per task and understand why long runs get expensive fast.
Each step re-reads everything so far
At each step the harness sends the model the whole conversation so far, including every earlier tool result. So input tokens per step grow, and the total over a run grows faster than the number of steps (roughly with the square of the steps when each step adds a similar amount). Output tokens are fewer but usually cost more per token. Levers: prompt caching (many providers bill repeated prefixes such as instructions and earlier turns at a lower rate), concise tool outputs, compaction, smaller models for easy sub-steps (searching, summarising) and step and cost caps. Track cost per completed task (including failed attempts), not cost per run, and compare it with the engineer time saved. The prices in the example are invented for illustration.
Cost growth with steps, run
I ran this with plain Python 3 (standard library only), using a throwaway project created in a temporary folder. With invented prices, a 5-step run uses 27,000 input tokens and costs about $0.11, while a 40-step run uses over a million input tokens and costs about $3.41: 8 times the steps, about 30 times the cost. Caching 80% of the input cuts the 40-step cost to about $1.13.
# Each step re-sends the growing conversation, so input tokens grow faster than the step count.
base, per_step = 3000, 1200 # initial prompt tokens, and tokens added by every step (tool output + reply)
price_in, price_out = 3.0, 15.0 # example USD per million tokens, NOT real prices
out_per_step = 400
def run_cost(steps, cached_fraction=0.0, cache_price=0.1):
total_in = sum(base + per_step * i for i in range(steps))
cached = total_in * cached_fraction
cost_in = ((total_in - cached) * price_in + cached * price_in * cache_price) / 1e6
return total_in, cost_in + steps * out_per_step * price_out / 1e6
for steps in (5, 10, 20, 40):
t, c = run_cost(steps)
print(f"{steps:2d} steps: input tokens {t:8,d} cost ${c:.3f} (with 80% of input cached: ${run_cost(steps, 0.8)[1]:.3f})")
Output:
5 steps: input tokens 27,000 cost $0.111 (with 80% of input cached: $0.053) 10 steps: input tokens 84,000 cost $0.312 (with 80% of input cached: $0.131) 20 steps: input tokens 288,000 cost $0.984 (with 80% of input cached: $0.362) 40 steps: input tokens 1,056,000 cost $3.408 (with 80% of input cached: $1.127)
Track cost per completed task
Include failed attempts in the cost of the tasks that eventually succeed.
Quick check: Why does cost rise faster than the number of steps?
- Each step re-sends a longer conversation as input
- Tokens get more expensive during the run
- The model charges per file
- It does not
Answer
Each step re-sends a longer conversation as input — Context growth makes later steps more expensive than earlier ones.