# Token Economics — AI Safety, Evaluation and Cost Control

Source: https://www.geekswithgeeks.com/en/ai-safety/cost-token-economics

> Estimate per-request and monthly cost from input tokens, output tokens and prices.

## Input and output are priced differently

Providers charge per million tokens, usually **more for output than for input**. Cost per request = input tokens × input price + output tokens × output price. Multiply by your monthly request count for a budget. Often the largest cost driver is **input size** (long prompts, long history, many retrieved documents) multiplied by traffic, so trimming context can matter more than switching providers. Prices change often; use your provider's current price page.

## Where the money goes and how to save it

Cost depends on tokens, model choice and number of calls. Routing, caching and trimming reduce all three.

![Four levers: route, cache, trim, cap.](assets/figures/ai-safety/section-7-map.svg) — Figure 7.1 — Route, cache, trim and cap.

## A monthly estimate, run

I ran this with illustrative prices of 3 and 15 per million input and output tokens. 100,000 requests a month at 1,500 input and 300 output tokens each cost 900.

```python
def monthly(requests, in_tok, out_tok, p_in, p_out):
    per_request = in_tok / 1e6 * p_in + out_tok / 1e6 * p_out
    return round(requests * per_request, 2)

print(monthly(100_000, 1500, 300, 3.0, 15.0))
```

Output:

```
900.0
```

## Cost per successful task

Cheap calls that fail and need retries or human fixes cost more than a pricier call that works. Track cost per successfully completed task, not only cost per token.

**Quiz:** Which usually drives LLM cost most in a chat product?

- [ ] The colour of the UI
- [x] Large inputs repeated across many requests
- [ ] The font size
- [ ] The number of buttons

*Answer:* Large inputs repeated across many requests. Long prompts and histories are paid for again on every call.
