Lesson 17 / 25

A/B Tests and Regression Checks

Compare versions on live traffic carefully and run offline evals in CI on every change.

Offline first, then online

Run your evaluation set automatically (regression check) whenever a prompt, model, tool or retrieval setting changes, and block the change if key scores drop. For user-facing results, run an A/B test: send a share of traffic to the new version and compare a metric such as thumbs-up rate or task completion. Use a significance test so random noise is not mistaken for improvement. For two proportions, a z-score near or above about 1.96 indicates a difference unlikely to be chance (at the usual 5% level).

Two-proportion z-score, run

I ran this. 120/1000 vs 150/1000 gives z = 1.96, right at the edge of significance. The same rates on only 100 users each (12 vs 15) give z = 0.62, which is not convincing.

import math

def z2(a, na, b, nb):
    p1, p2 = a / na, b / nb
    p = (a + b) / (na + nb)
    se = math.sqrt(p * (1 - p) * (1 / na + 1 / nb))
    return round((p2 - p1) / se, 2)

print(z2(120, 1000, 150, 1000), z2(12, 100, 15, 100))

Output:

1.96 0.62

Decide the metric and stopping rule first

Choosing the metric after seeing results, or stopping the test the moment it looks good, inflates false wins. Write down what you will measure and how long you will run before you start.

Quick check: Why block a change when regression-check scores drop?

  • Dropping scores are always noise
  • Scores never matter
  • The change may have made the system worse in ways you did not intend
  • It speeds up deployment
Answer

The change may have made the system worse in ways you did not intend — Automatic checks catch silent quality loss before users do.