Lesson 23 / 25
Budgets, Quotas and FinOps for AI
Set spend limits, allocate costs to features and alert before budgets are exceeded.
Make spend visible and bounded
Tag each request with the feature, team and environment so you can see who spends what. Set hard limits (provider-side spend caps, per-key quotas, per-user token limits) and soft alerts at, say, 50%, 80% and 100% of the monthly budget. Use separate API keys for development, testing and production so a test script cannot drain the production budget. Review the bill regularly and ask whether each large cost is buying value.
A budget policy
Write it down so alerts and limits have owners. Numbers here are illustrative.
Monthly AI budget 1,000 USD (support assistant 600, search 250, internal 150)
Soft alerts 50% / 80% / 100% -> Slack #ai-costs + feature owner
Hard limits provider spend cap at 110%; per-user 200k tokens/day
Keys dev, staging, prod are separate; prod key not in CI
Monthly review cost per resolved ticket; top 5 costly prompts; routing hit-rateQuick check: Why use separate API keys for development and production?
- So a test script cannot drain the production budget, and spend is attributable
- Keys must be unique by law
- It makes the model faster
- It removes the need for budgets
Answer
So a test script cannot drain the production budget, and spend is attributable — Separate keys isolate environments and make costs traceable.