Lesson 19 / 25

Latency Percentiles and Alerts

Track p50 and p95, not only the average, and alert on meaningful thresholds.

Averages hide slow requests

A few very slow requests barely move the average but ruin the experience for those users. Track percentiles: p50 (the typical request) and p95 or p99 (the slow tail). Set alerts for sustained problems, such as p95 latency above a limit, error rate above 2%, moderation blocks suddenly doubling, thumbs-down rate rising, or daily spend above budget. Alert on trends over a window, not single spikes, to avoid constant false alarms.

Mean vs percentiles, run

I ran this on ten latencies in seconds. The mean is 2.02 s, the median (p50) is 1.25 s, and p95 is 5.85 s because of one 9-second outlier. The average understates how slow the slowest requests are.

import statistics
lat = [0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 2.0, 9.0]

def pct(x, q):
    x = sorted(x)
    k = (len(x) - 1) * q
    f = int(k)
    c = min(f + 1, len(x) - 1)
    return round(x[f] + (x[c] - x[f]) * (k - f), 2)

print(round(statistics.mean(lat), 2), pct(lat, 0.5), pct(lat, 0.95))

Output:

2.02 1.25 5.85

Stream responses

Streaming tokens as they are generated makes slow answers feel faster, but measure both time to first token and total time, because they affect users differently.

Quick check: Why track p95 latency in addition to the average?

  • It reduces the bill
  • p95 is always smaller than the mean
  • Averages are illegal
  • The slow tail affects real users but barely moves the average
Answer

The slow tail affects real users but barely moves the average — Percentiles expose the worst experiences that an average smooths over.