Lesson 21 / 25
Model Routing and Cascades
Send easy requests to a small, cheap model and hard ones to a stronger model.
Right-size the model
Not every request needs the largest model. A router sends simple, well-defined requests (classification, extraction, FAQ-style questions) to a small model and only escalates difficult ones to a stronger model. A cascade tries the small model first and escalates when its answer fails a check or has low confidence. Savings depend on how much traffic is easy and how often the router is wrong, so measure quality on both paths with your evaluation set.
Savings from routing, run
I ran this with illustrative prices. A small model at 0.25 / 1.25 costs 75 for the same traffic; sending 70% of requests there and 30% to the large model costs 322.5 instead of 900, a saving of about 64%, provided quality on the easy 70% stays acceptable.
def monthly(requests, in_tok, out_tok, p_in, p_out):
return round(requests * (in_tok / 1e6 * p_in + out_tok / 1e6 * p_out), 2)
base = monthly(100_000, 1500, 300, 3.0, 15.0) # large model for everything
small = monthly(100_000, 1500, 300, 0.25, 1.25) # small model for everything
routed = round(0.7 * small + 0.3 * base, 2) # 70% easy -> small model
print(base, small, routed, round(1 - routed / base, 3))
Output:
900.0 75.0 322.5 0.642
Check quality on the cheap path
Savings mean nothing if the small model quietly lowers quality. Run your evaluation set on the routed system, including a per-slice report, before and after switching.
Quick check: What is the main risk of routing requests to a cheaper model?
- The bill always rises
- Quality may drop on requests the router misjudged as easy
- Small models cannot read text
- Routing is illegal
Answer
Quality may drop on requests the router misjudged as easy — Savings depend on the router being right, so quality must be measured on the cheap path.