# Drift: When the World or the Model Changes — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/r-drift

> Detect shifts in inputs and quality before users complain.

## Yesterday's evaluation may not describe today

A system that scored well at launch can degrade without any code change: users ask about **new topics** after a product launch, a **provider updates the model**, documents in the index go **stale**, seasonal traffic shifts the mix, or the language mix changes. Detect it with: **input monitoring** (distribution of topics, languages, lengths, compared with a baseline, for example using the population stability index, where PSI under 0.1 is stable and over 0.25 is a large shift); **quality sampling** (score a random sample of live traffic regularly, with the judge plus human spot checks); **output monitoring** (refusal rate, validation failures, length, cost per request); and **user feedback** (thumbs, edits, escalations). Run the full evaluation again on a schedule and after any provider notice, and feed newly seen failure types back into the test set.

## Population stability index, run

I ran this with plain Python 3 (standard library only). Comparing this week's topic mix with the launch baseline gives a PSI of 0.002 (stable). After a product launch changes what people ask, the PSI is 0.559, a large shift that should trigger re-running the evaluations and reviewing prompts and documents.

```python
import math

def psi(expected, actual):
    """Population Stability Index between two distributions given as proportions per bucket."""
    return sum((a - e) * math.log(a / e) for e, a in zip(expected, actual))

baseline = [0.50, 0.30, 0.15, 0.05]       # share of queries by topic when the app launched
this_week = [0.48, 0.31, 0.16, 0.05]
new_topic = [0.30, 0.25, 0.15, 0.30]      # a new product launch changed what people ask
for name, dist in (("this week", this_week), ("after product launch", new_topic)):
    v = psi(baseline, dist)
    verdict = "stable" if v < 0.1 else "moderate shift" if v < 0.25 else "large shift: re-run evals, review prompts"
    print(f"{name:21} PSI {v:.3f} -> {verdict}")

```

Output:

```
this week             PSI 0.002 -> stable
after product launch  PSI 0.559 -> large shift: re-run evals, review prompts
```

## Feed failures back into the test set

Real failures are the best new test cases.

**Quiz:** A system's quality drops although no code changed. Which is a plausible cause?

- [ ] The code compiled differently
- [x] The provider updated the model, or users now ask about new topics
- [ ] Python forgot how to run
- [ ] It cannot happen

*Answer:* The provider updated the model, or users now ask about new topics. Inputs, data and the model can all change under you.
