# Cost, Latency and Caching — Large Language Models

Source: https://www.geekswithgeeks.com/en/llms/u-cost

> Estimate API cost from tokens and reduce it.

## You pay per token

Hosted APIs charge per million **input tokens** and per million **output tokens**, with output usually priced higher. Latency has two parts: **time to first token** (reading the prompt) and **generation speed** (tokens per second), and long outputs take longer. To cut cost and delay: shorten prompts, retrieve only relevant chunks, cap output length, use a **smaller model** for easy tasks, **cache** repeated prompt prefixes where the provider supports it, and **batch** non-urgent work. The prices below are example numbers, not any provider's real price list.

## A monthly cost estimate, run

I ran this plain-Python (standard library only) example. Ten thousand calls with a 2,000-token prompt and 500-token answer cost $135 at the example prices; caching the prompt at 10% of its price would cut that to $81.

```python
price_in, price_out = 3.0, 15.0     # example prices, USD per million tokens
prompt_tokens, output_tokens, calls = 2000, 500, 10000
cost = calls * (prompt_tokens * price_in + output_tokens * price_out) / 1e6
print("monthly cost: $", cost)
cached = calls * (prompt_tokens * price_in * 0.1 + output_tokens * price_out) / 1e6
print("if the 2000-token prompt were cached at 10% price: $", cached)

```

Output:

```
monthly cost: $ 135.0
if the 2000-token prompt were cached at 10% price: $ 81.0
```

**Quiz:** Which change usually lowers cost the most for easy classification tasks?

- [ ] Using the largest model with a long prompt
- [x] Using a smaller model with a short prompt
- [ ] Raising the temperature
- [ ] Adding more examples to every call

*Answer:* Using a smaller model with a short prompt. Match model size and prompt length to the difficulty of the task.
