Lesson 9 / 25

Counting Tokens

Read usage from responses and estimate size before you send.

Measure, do not guess

Every Claude API response includes a usage object with input_tokens and output_tokens; with prompt caching it also reports cache creation and cache read tokens. Add these up each turn to feed your Budget. To check a prompt before sending, use the provider's token-counting endpoint; as a rough guide English text is about 4 characters per token, and Hindi or code often uses more tokens per character.

Where the budget goes

A context window is shared by the system prompt, tools, history and the reply. Reserve space for each.

Three shares: fixed prompt, growing history, reply reserve.
Figure 3.1 — Fixed prompt, growing history and reply reserve.

Charging the budget from usage

response is the API reply. Input and output tokens are billed at different prices, so keep them separate.

u = response.usage
in_tokens, out_tokens = u.input_tokens, u.output_tokens
budget.charge(in_tokens + out_tokens)
metrics["input"] += in_tokens
metrics["output"] += out_tokens

A rough size estimate

A quick estimate of 4 characters per token is fine for guard rails but not for billing. 1,200 characters gives about 300 tokens.

def rough_tokens(text: str) -> int:
    return len(text) // 4

print(rough_tokens("hello world " * 100))

Output:

300

Quick check: Where do you find the exact tokens a call used?

  • The usage object in the response
  • The length of the prompt in characters
  • The model name
  • The HTTP status code
Answer

The usage object in the response — The API reports exact input and output token counts for each response.