Lesson 20 / 27
Cost, Latency and Caching
Estimate API cost from tokens and reduce it.
You pay per token
Hosted APIs charge per million input tokens and per million output tokens, with output usually priced higher. Latency has two parts: time to first token (reading the prompt) and generation speed (tokens per second), and long outputs take longer. To cut cost and delay: shorten prompts, retrieve only relevant chunks, cap output length, use a smaller model for easy tasks, cache repeated prompt prefixes where the provider supports it, and batch non-urgent work. The prices below are example numbers, not any provider's real price list.
A monthly cost estimate, run
I ran this plain-Python (standard library only) example. Ten thousand calls with a 2,000-token prompt and 500-token answer cost $135 at the example prices; caching the prompt at 10% of its price would cut that to $81.
price_in, price_out = 3.0, 15.0 # example prices, USD per million tokens
prompt_tokens, output_tokens, calls = 2000, 500, 10000
cost = calls * (prompt_tokens * price_in + output_tokens * price_out) / 1e6
print("monthly cost: $", cost)
cached = calls * (prompt_tokens * price_in * 0.1 + output_tokens * price_out) / 1e6
print("if the 2000-token prompt were cached at 10% price: $", cached)
Output:
monthly cost: $ 135.0 if the 2000-token prompt were cached at 10% price: $ 81.0
Quick check: Which change usually lowers cost the most for easy classification tasks?
- Using the largest model with a long prompt
- Using a smaller model with a short prompt
- Raising the temperature
- Adding more examples to every call
Answer
Using a smaller model with a short prompt — Match model size and prompt length to the difficulty of the task.