Lesson 20 / 29
Drift: When the World or the Model Changes
Detect shifts in inputs and quality before users complain.
Yesterday's evaluation may not describe today
A system that scored well at launch can degrade without any code change: users ask about new topics after a product launch, a provider updates the model, documents in the index go stale, seasonal traffic shifts the mix, or the language mix changes. Detect it with: input monitoring (distribution of topics, languages, lengths, compared with a baseline, for example using the population stability index, where PSI under 0.1 is stable and over 0.25 is a large shift); quality sampling (score a random sample of live traffic regularly, with the judge plus human spot checks); output monitoring (refusal rate, validation failures, length, cost per request); and user feedback (thumbs, edits, escalations). Run the full evaluation again on a schedule and after any provider notice, and feed newly seen failure types back into the test set.
Population stability index, run
I ran this with plain Python 3 (standard library only). Comparing this week's topic mix with the launch baseline gives a PSI of 0.002 (stable). After a product launch changes what people ask, the PSI is 0.559, a large shift that should trigger re-running the evaluations and reviewing prompts and documents.
import math
def psi(expected, actual):
"""Population Stability Index between two distributions given as proportions per bucket."""
return sum((a - e) * math.log(a / e) for e, a in zip(expected, actual))
baseline = [0.50, 0.30, 0.15, 0.05] # share of queries by topic when the app launched
this_week = [0.48, 0.31, 0.16, 0.05]
new_topic = [0.30, 0.25, 0.15, 0.30] # a new product launch changed what people ask
for name, dist in (("this week", this_week), ("after product launch", new_topic)):
v = psi(baseline, dist)
verdict = "stable" if v < 0.1 else "moderate shift" if v < 0.25 else "large shift: re-run evals, review prompts"
print(f"{name:21} PSI {v:.3f} -> {verdict}")
Output:
this week PSI 0.002 -> stable after product launch PSI 0.559 -> large shift: re-run evals, review prompts
Feed failures back into the test set
Real failures are the best new test cases.
Quick check: A system's quality drops although no code changed. Which is a plausible cause?
- The code compiled differently
- The provider updated the model, or users now ask about new topics
- Python forgot how to run
- It cannot happen
Answer
The provider updated the model, or users now ask about new topics — Inputs, data and the model can all change under you.