Lesson 3 / 25
Why Cost Grows Faster Than Steps
Compute total input tokens across a loop and see why long loops get expensive quickly.
You resend the whole history
The API is stateless: every turn resends all earlier messages. If each turn adds about 1,000 tokens, turn 1 sends 1,000, turn 2 sends 2,000, turn 3 sends 3,000, and so on. Total input over n turns is roughly 1,000 × (1 + 2 + … + n), which grows with the square of n, not linearly.
Ten turns, run it
Ten turns that each add 1,000 tokens send 55,000 input tokens in total, not 10,000. The cost line uses illustrative prices of 3 and 15 per million tokens; use your provider's real prices.
per_turn, total = 1000, 0
for turn in range(1, 11):
total += per_turn * turn # history so far is resent
print(total)
def cost(inp, out, in_price, out_price):
return inp / 1_000_000 * in_price + out / 1_000_000 * out_price
print(round(cost(55_000, 3_000, 3.0, 15.0), 4))
Output:
55000 0.21
Rereading the whole book
Imagine that before writing each new page you must reread every earlier page. The first pages are quick, but by page 50 you spend most of your time rereading. That is an agent loop without trimming or caching.
Quick check: Each turn adds 1,000 tokens. About how many input tokens do 10 turns send in total?
- 10,000
- 100,000
- 55,000
- 1,000
Answer
55,000 — Because history is resent, the total is 1,000 × (1 + … + 10) = 55,000.