Lesson 20 / 25
Token Economics
Estimate per-request and monthly cost from input tokens, output tokens and prices.
Input and output are priced differently
Providers charge per million tokens, usually more for output than for input. Cost per request = input tokens × input price + output tokens × output price. Multiply by your monthly request count for a budget. Often the largest cost driver is input size (long prompts, long history, many retrieved documents) multiplied by traffic, so trimming context can matter more than switching providers. Prices change often; use your provider's current price page.
Where the money goes and how to save it
Cost depends on tokens, model choice and number of calls. Routing, caching and trimming reduce all three.
A monthly estimate, run
I ran this with illustrative prices of 3 and 15 per million input and output tokens. 100,000 requests a month at 1,500 input and 300 output tokens each cost 900.
def monthly(requests, in_tok, out_tok, p_in, p_out):
per_request = in_tok / 1e6 * p_in + out_tok / 1e6 * p_out
return round(requests * per_request, 2)
print(monthly(100_000, 1500, 300, 3.0, 15.0))
Output:
900.0
Cost per successful task
Cheap calls that fail and need retries or human fixes cost more than a pricier call that works. Track cost per successfully completed task, not only cost per token.
Quick check: Which usually drives LLM cost most in a chat product?
- The colour of the UI
- Large inputs repeated across many requests
- The font size
- The number of buttons
Answer
Large inputs repeated across many requests — Long prompts and histories are paid for again on every call.