Lesson 20 / 25

Token Economics

Estimate per-request and monthly cost from input tokens, output tokens and prices.

Input and output are priced differently

Providers charge per million tokens, usually more for output than for input. Cost per request = input tokens × input price + output tokens × output price. Multiply by your monthly request count for a budget. Often the largest cost driver is input size (long prompts, long history, many retrieved documents) multiplied by traffic, so trimming context can matter more than switching providers. Prices change often; use your provider's current price page.

Where the money goes and how to save it

Cost depends on tokens, model choice and number of calls. Routing, caching and trimming reduce all three.

Four levers: route, cache, trim, cap.
Figure 7.1 — Route, cache, trim and cap.

A monthly estimate, run

I ran this with illustrative prices of 3 and 15 per million input and output tokens. 100,000 requests a month at 1,500 input and 300 output tokens each cost 900.

def monthly(requests, in_tok, out_tok, p_in, p_out):
    per_request = in_tok / 1e6 * p_in + out_tok / 1e6 * p_out
    return round(requests * per_request, 2)

print(monthly(100_000, 1500, 300, 3.0, 15.0))

Output:

900.0

Cost per successful task

Cheap calls that fail and need retries or human fixes cost more than a pricier call that works. Track cost per successfully completed task, not only cost per token.

Quick check: Which usually drives LLM cost most in a chat product?

  • The colour of the UI
  • Large inputs repeated across many requests
  • The font size
  • The number of buttons
Answer

Large inputs repeated across many requests — Long prompts and histories are paid for again on every call.