Lesson 26 / 29
Cost, Latency and Prompt Length
Trim prompts and choose models to keep systems affordable and fast.
Every token is paid for, every call
Providers charge per input and output token, so a long fixed prefix is paid for on every call. Reduce cost and delay by: cutting redundant instructions and unneeded examples, shortening retrieved context, capping max output tokens, using a smaller model for easy tasks, caching repeated prefixes (prompt caching) where available, and batching non-urgent work. Always re-run the test set after shortening a prompt to confirm quality held. The numbers below use example prices, not any provider's real price list.
What a shorter prompt saves, run
I ran this plain-Python (standard library only) example. At the example prices, cutting the prompt from 1,800 to 600 tokens saves $360 per 100,000 calls (from $840 to $480).
def cost(prompt_tokens, output_tokens, price_in=3.0, price_out=15.0, calls=1):
return calls * (prompt_tokens * price_in + output_tokens * price_out) / 1e6
long_p, short_p = 1800, 600
print("long prompt per 100k calls: $", round(cost(long_p, 200, calls=100_000), 2))
print("short prompt per 100k calls: $", round(cost(short_p, 200, calls=100_000), 2))
print("saving: $", round(cost(long_p, 200, calls=100_000) - cost(short_p, 200, calls=100_000), 2))
Output:
long prompt per 100k calls: $ 840.0 short prompt per 100k calls: $ 480.0 saving: $ 360.0
Quick check: Why does a long fixed prompt prefix matter for cost?
- It is paid for on every single call
- It is free
- It is paid once per year
- It reduces output tokens
Answer
It is paid for on every single call — Fixed overhead multiplies with call volume.