Lesson 9 / 25
Counting Tokens
Read usage from responses and estimate size before you send.
Measure, do not guess
Every Claude API response includes a usage object with input_tokens and output_tokens; with prompt caching it also reports cache creation and cache read tokens. Add these up each turn to feed your Budget. To check a prompt before sending, use the provider's token-counting endpoint; as a rough guide English text is about 4 characters per token, and Hindi or code often uses more tokens per character.
Where the budget goes
A context window is shared by the system prompt, tools, history and the reply. Reserve space for each.
Charging the budget from usage
response is the API reply. Input and output tokens are billed at different prices, so keep them separate.
u = response.usage
in_tokens, out_tokens = u.input_tokens, u.output_tokens
budget.charge(in_tokens + out_tokens)
metrics["input"] += in_tokens
metrics["output"] += out_tokensA rough size estimate
A quick estimate of 4 characters per token is fine for guard rails but not for billing. 1,200 characters gives about 300 tokens.
def rough_tokens(text: str) -> int:
return len(text) // 4
print(rough_tokens("hello world " * 100))
Output:
300
Quick check: Where do you find the exact tokens a call used?
- The usage object in the response
- The length of the prompt in characters
- The model name
- The HTTP status code
Answer
The usage object in the response — The API reports exact input and output token counts for each response.