# Model Routing and Cascades — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/p-route

> Send easy requests to cheap models and hard ones to strong models.

## Not every question needs the biggest model

Strong models cost many times more than small ones, but many requests are easy. A **cascade** tries a cheap model first and **escalates** to a stronger one when a **confidence signal** says the cheap answer may be wrong (low self-reported confidence, failed validation, an unsure classifier, a disagreement between samples). A **router** instead predicts difficulty up front and picks the model. Both trade a little quality for a large cost saving, and both depend on how good the signal is. Evaluate the system **end to end**: measure overall quality and cost against the strong-only baseline on your test set, and look at quality per category, since the cheap path may fail on exactly the hard categories you care about. Keep fallbacks for outages: if the main model is down, fail over to another.

## A cascade versus strong-only (simulation), run

I ran this with plain Python 3 (standard library only). This is a simulation with invented numbers and seeded random draws, so it shows the shape of the effect, not measurements of any product. With 10,000 simulated queries, the cascade reaches quality 0.902 at 43% of the strong model's cost, against 0.971 for the strong model alone. That is a real trade: about 7 points of quality for more than half the cost. Whether it is acceptable depends on your quality target and on which categories lose.

```python
import random
rng = random.Random(42)

N = 10000
# each query has a difficulty in [0,1]; the cheap model solves easy ones, the strong model solves almost all
queries = [rng.random() for _ in range(N)]
cheap_ok  = lambda d: rng.random() < (0.95 if d < 0.6 else 0.35)
strong_ok = lambda d: rng.random() < 0.97
cheap_cost, strong_cost = 1.0, 12.0                         # relative units
def confident(d): return d < 0.55 or rng.random() < 0.25     # imperfect confidence signal from the cheap model

only_strong = sum(strong_ok(d) for d in queries) / N
ok = cost = 0
for d in queries:
    cost += cheap_cost
    if confident(d):
        ok += cheap_ok(d)
    else:
        cost += strong_cost; ok += strong_ok(d)
print(f"strong model only : quality {only_strong:.3f}, cost {strong_cost * N:,.0f}")
print(f"cascade           : quality {ok / N:.3f}, cost {cost:,.0f}  ({cost / (strong_cost * N):.0%} of the strong-only cost)")

```

Output:

```
strong model only : quality 0.971, cost 120,000
cascade           : quality 0.902, cost 51,388  (43% of the strong-only cost)
```

## Keep a fallback model

If the main model is down, fail over to another so users are still served.

**Quiz:** How should a cascade be evaluated?

- [x] End to end, comparing quality and cost with a strong-model baseline, including per category
- [ ] Only by its cost
- [ ] Only by the cheap model's score
- [ ] It cannot be evaluated

*Answer:* End to end, comparing quality and cost with a strong-model baseline, including per category. A cheap path that fails on hard categories can hide a serious quality loss.
