Lesson 23 / 25

Budgets, Quotas and FinOps for AI

Set spend limits, allocate costs to features and alert before budgets are exceeded.

Make spend visible and bounded

Tag each request with the feature, team and environment so you can see who spends what. Set hard limits (provider-side spend caps, per-key quotas, per-user token limits) and soft alerts at, say, 50%, 80% and 100% of the monthly budget. Use separate API keys for development, testing and production so a test script cannot drain the production budget. Review the bill regularly and ask whether each large cost is buying value.

A budget policy

Write it down so alerts and limits have owners. Numbers here are illustrative.

Monthly AI budget      1,000 USD  (support assistant 600, search 250, internal 150)
Soft alerts             50% / 80% / 100%  -> Slack #ai-costs + feature owner
Hard limits             provider spend cap at 110%; per-user 200k tokens/day
Keys                    dev, staging, prod are separate; prod key not in CI
Monthly review          cost per resolved ticket; top 5 costly prompts; routing hit-rate

Quick check: Why use separate API keys for development and production?

  • So a test script cannot drain the production budget, and spend is attributable
  • Keys must be unique by law
  • It makes the model faster
  • It removes the need for budgets
Answer

So a test script cannot drain the production budget, and spend is attributable — Separate keys isolate environments and make costs traceable.