Lesson 19 / 25
Latency Percentiles and Alerts
Track p50 and p95, not only the average, and alert on meaningful thresholds.
Averages hide slow requests
A few very slow requests barely move the average but ruin the experience for those users. Track percentiles: p50 (the typical request) and p95 or p99 (the slow tail). Set alerts for sustained problems, such as p95 latency above a limit, error rate above 2%, moderation blocks suddenly doubling, thumbs-down rate rising, or daily spend above budget. Alert on trends over a window, not single spikes, to avoid constant false alarms.
Mean vs percentiles, run
I ran this on ten latencies in seconds. The mean is 2.02 s, the median (p50) is 1.25 s, and p95 is 5.85 s because of one 9-second outlier. The average understates how slow the slowest requests are.
import statistics
lat = [0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 2.0, 9.0]
def pct(x, q):
x = sorted(x)
k = (len(x) - 1) * q
f = int(k)
c = min(f + 1, len(x) - 1)
return round(x[f] + (x[c] - x[f]) * (k - f), 2)
print(round(statistics.mean(lat), 2), pct(lat, 0.5), pct(lat, 0.95))
Output:
2.02 1.25 5.85
Stream responses
Streaming tokens as they are generated makes slow answers feel faster, but measure both time to first token and total time, because they affect users differently.
Quick check: Why track p95 latency in addition to the average?
- It reduces the bill
- p95 is always smaller than the mean
- Averages are illegal
- The slow tail affects real users but barely moves the average
Answer
The slow tail affects real users but barely moves the average — Percentiles expose the worst experiences that an average smooths over.