Lesson 20 / 29

Drift: When the World or the Model Changes

Detect shifts in inputs and quality before users complain.

Yesterday's evaluation may not describe today

A system that scored well at launch can degrade without any code change: users ask about new topics after a product launch, a provider updates the model, documents in the index go stale, seasonal traffic shifts the mix, or the language mix changes. Detect it with: input monitoring (distribution of topics, languages, lengths, compared with a baseline, for example using the population stability index, where PSI under 0.1 is stable and over 0.25 is a large shift); quality sampling (score a random sample of live traffic regularly, with the judge plus human spot checks); output monitoring (refusal rate, validation failures, length, cost per request); and user feedback (thumbs, edits, escalations). Run the full evaluation again on a schedule and after any provider notice, and feed newly seen failure types back into the test set.

Population stability index, run

I ran this with plain Python 3 (standard library only). Comparing this week's topic mix with the launch baseline gives a PSI of 0.002 (stable). After a product launch changes what people ask, the PSI is 0.559, a large shift that should trigger re-running the evaluations and reviewing prompts and documents.

import math

def psi(expected, actual):
    """Population Stability Index between two distributions given as proportions per bucket."""
    return sum((a - e) * math.log(a / e) for e, a in zip(expected, actual))

baseline = [0.50, 0.30, 0.15, 0.05]       # share of queries by topic when the app launched
this_week = [0.48, 0.31, 0.16, 0.05]
new_topic = [0.30, 0.25, 0.15, 0.30]      # a new product launch changed what people ask
for name, dist in (("this week", this_week), ("after product launch", new_topic)):
    v = psi(baseline, dist)
    verdict = "stable" if v < 0.1 else "moderate shift" if v < 0.25 else "large shift: re-run evals, review prompts"
    print(f"{name:21} PSI {v:.3f} -> {verdict}")

Output:

this week             PSI 0.002 -> stable
after product launch  PSI 0.559 -> large shift: re-run evals, review prompts

Feed failures back into the test set

Real failures are the best new test cases.

Quick check: A system's quality drops although no code changed. Which is a plausible cause?

  • The code compiled differently
  • The provider updated the model, or users now ask about new topics
  • Python forgot how to run
  • It cannot happen
Answer

The provider updated the model, or users now ask about new topics — Inputs, data and the model can all change under you.