# Latency Percentiles and Alerts — AI Safety, Evaluation and Cost Control

Source: https://www.geekswithgeeks.com/en/ai-safety/mon-latency-alerts

> Track p50 and p95, not only the average, and alert on meaningful thresholds.

## Averages hide slow requests

A few very slow requests barely move the average but ruin the experience for those users. Track **percentiles**: **p50** (the typical request) and **p95** or **p99** (the slow tail). Set **alerts** for sustained problems, such as p95 latency above a limit, error rate above 2%, moderation blocks suddenly doubling, thumbs-down rate rising, or daily spend above budget. Alert on trends over a window, not single spikes, to avoid constant false alarms.

## Mean vs percentiles, run

I ran this on ten latencies in seconds. The mean is 2.02 s, the median (p50) is 1.25 s, and p95 is 5.85 s because of one 9-second outlier. The average understates how slow the slowest requests are.

```python
import statistics
lat = [0.8, 0.9, 1.0, 1.1, 1.2, 1.3, 1.4, 1.5, 2.0, 9.0]

def pct(x, q):
    x = sorted(x)
    k = (len(x) - 1) * q
    f = int(k)
    c = min(f + 1, len(x) - 1)
    return round(x[f] + (x[c] - x[f]) * (k - f), 2)

print(round(statistics.mean(lat), 2), pct(lat, 0.5), pct(lat, 0.95))
```

Output:

```
2.02 1.25 5.85
```

## Stream responses

Streaming tokens as they are generated makes slow answers feel faster, but measure both time to first token and total time, because they affect users differently.

**Quiz:** Why track p95 latency in addition to the average?

- [ ] It reduces the bill
- [ ] p95 is always smaller than the mean
- [ ] Averages are illegal
- [x] The slow tail affects real users but barely moves the average

*Answer:* The slow tail affects real users but barely moves the average. Percentiles expose the worst experiences that an average smooths over.
