# Tokens, Pricing and Estimating Cost — Claude API / OpenAI API Basics

Source: https://www.geekswithgeeks.com/en/llm-apis/c-tokens

> Turn usage numbers into money before you ship.

## Input and output are priced separately

Providers price per million **input tokens** and per million **output tokens**, with output usually costing several times more, and bigger or newer models costing more than smaller ones. Some features (cached prompt prefixes, batch processing, images, tool definitions, long context) have their own rates. So a call's cost is `(input_tokens x input_price + output_tokens x output_price) / 1,000,000`. Estimate before launch: average input and output tokens per request times expected requests per month. Remember that **tool definitions, system prompts and conversation history count as input tokens on every request**. Prices change, so read the provider's current pricing page; the figures below are **made-up examples**, not real prices.

## Know what each call costs

Cost is tokens times price; trim prompts, cache prefixes, pick the right model and track usage.

![Three levers: count, cache, choose.](assets/figures/llm-apis/section-6-map.svg) — Figure 6.1 — Count, cache and choose.

## A cost calculator, run

I ran this plain-Python (standard library only) example. With invented prices, the same 1,200-in / 150-out call costs about $0.0005 on the small model and $0.0059 on the large one; a long 8,000-token prompt on the large model costs $0.033. The last line shows how a tiny per-call cost grows to $4.88 over 10,000 calls.

```python
PRICES = {"small-model": (0.25, 1.25), "large-model": (3.00, 15.00)}      # example USD per million tokens, NOT real prices

def cost(model, input_tokens, output_tokens):
    pin, pout = PRICES[model]
    return (input_tokens * pin + output_tokens * pout) / 1_000_000

calls = [("small-model", 1200, 150), ("large-model", 1200, 150), ("large-model", 8000, 600)]
total = 0.0
for model, i, o in calls:
    c = cost(model, i, o); total += c
    print(f"{model:12} in={i:5} out={o:4} -> ${c:.6f}")
print("total: $%.6f" % total)
print("10,000 calls like the first one: $%.2f" % (10_000 * cost("small-model", 1200, 150)))

```

Output:

```
small-model  in= 1200 out= 150 -> $0.000487
large-model  in= 1200 out= 150 -> $0.005850
large-model  in= 8000 out= 600 -> $0.033000
total: $0.039338
10,000 calls like the first one: $4.88
```

## Count tokens before sending

Use the provider's token counting to size prompts and reject oversized inputs early.

**Quiz:** What counts as input tokens on every request?

- [ ] Only the reply
- [ ] Only the newest user word
- [ ] Nothing
- [x] System prompt, tool definitions and conversation history

*Answer:* System prompt, tool definitions and conversation history. Everything you send is billed, including repeated context.
