Lesson 20 / 27
Tracking Usage and Setting Budgets
Measure per user, feature and model, and stop runaway spend.
What you do not measure will surprise you
Every response carries usage. Record it per request together with the user or tenant, feature, model and prompt version, so you can answer "which feature costs most?" and "which customer is expensive?". Set budgets and alerts in the provider console and in your own code (a per-user daily token limit, a per-request cap, a monthly ceiling that disables non-essential features). Add guardrails against abuse and bugs: cap max_tokens, cap loop steps in agents, limit input size, rate-limit users. Review the most expensive requests regularly; they often reveal a prompt that grew too long or a loop that never stops.
A budget guard, run
I ran this plain-Python (standard library only) example. After three calls the tracker has used 5,700 of a 6,000-token budget. A 1,000-token request would exceed it and is refused; a 300-token request fits exactly and is allowed.
class UsageTracker:
def __init__(self, budget_tokens):
self.budget, self.used, self.calls = budget_tokens, 0, 0
def record(self, input_tokens, output_tokens):
self.used += input_tokens + output_tokens; self.calls += 1
def allow_next(self, estimate):
return self.used + estimate <= self.budget
t = UsageTracker(budget_tokens=6000)
for inp, out in [(1200, 300), (1500, 400), (1800, 500)]:
t.record(inp, out)
print("used", t.used, "of", t.budget, "after", t.calls, "calls")
print("allow a 1000-token call?", t.allow_next(1000))
print("allow a 300-token call? ", t.allow_next(300))
Output:
used 5700 of 6000 after 3 calls allow a 1000-token call? False allow a 300-token call? True
Quick check: Why record usage with feature and user labels?
- To make the model faster
- To find which features and customers drive cost
- To avoid keys
- Because the API requires it
Answer
To find which features and customers drive cost — Attribution turns a single bill into actionable information.