# What a Run Costs: Context Growth and Caching — Coding Agents & AI-Assisted Development

Source: https://www.geekswithgeeks.com/en/coding-agents/m-cost

> Estimate cost per task and understand why long runs get expensive fast.

## Each step re-reads everything so far

At each step the harness sends the model the whole conversation so far, including every earlier tool result. So input tokens per step **grow**, and the total over a run grows faster than the number of steps (roughly with the square of the steps when each step adds a similar amount). Output tokens are fewer but usually cost more per token. Levers: **prompt caching** (many providers bill repeated prefixes such as instructions and earlier turns at a lower rate), **concise tool outputs**, **compaction**, **smaller models for easy sub-steps** (searching, summarising) and **step and cost caps**. Track **cost per completed task** (including failed attempts), not cost per run, and compare it with the engineer time saved. The prices in the example are invented for illustration.

## Cost growth with steps, run

I ran this with plain Python 3 (standard library only), using a throwaway project created in a temporary folder. With invented prices, a 5-step run uses 27,000 input tokens and costs about $0.11, while a 40-step run uses over a million input tokens and costs about $3.41: 8 times the steps, about 30 times the cost. Caching 80% of the input cuts the 40-step cost to about $1.13.

```python
# Each step re-sends the growing conversation, so input tokens grow faster than the step count.
base, per_step = 3000, 1200          # initial prompt tokens, and tokens added by every step (tool output + reply)
price_in, price_out = 3.0, 15.0      # example USD per million tokens, NOT real prices
out_per_step = 400

def run_cost(steps, cached_fraction=0.0, cache_price=0.1):
    total_in = sum(base + per_step * i for i in range(steps))
    cached = total_in * cached_fraction
    cost_in = ((total_in - cached) * price_in + cached * price_in * cache_price) / 1e6
    return total_in, cost_in + steps * out_per_step * price_out / 1e6

for steps in (5, 10, 20, 40):
    t, c = run_cost(steps)
    print(f"{steps:2d} steps: input tokens {t:8,d}  cost ${c:.3f}   (with 80% of input cached: ${run_cost(steps, 0.8)[1]:.3f})")

```

Output:

```
 5 steps: input tokens   27,000  cost $0.111   (with 80% of input cached: $0.053)
10 steps: input tokens   84,000  cost $0.312   (with 80% of input cached: $0.131)
20 steps: input tokens  288,000  cost $0.984   (with 80% of input cached: $0.362)
40 steps: input tokens 1,056,000  cost $3.408   (with 80% of input cached: $1.127)
```

## Track cost per completed task

Include failed attempts in the cost of the tasks that eventually succeed.

**Quiz:** Why does cost rise faster than the number of steps?

- [x] Each step re-sends a longer conversation as input
- [ ] Tokens get more expensive during the run
- [ ] The model charges per file
- [ ] It does not

*Answer:* Each step re-sends a longer conversation as input. Context growth makes later steps more expensive than earlier ones.
