Lesson 16 / 29
Tail Latency, Timeouts and Hedged Requests
Manage the slow 1% that users remember.
Averages hide the slow requests
LLM latency has a long tail: most calls are quick, a few take many times longer (queueing, long outputs, provider hiccups). Users judge a product by its slow moments, so track percentiles (p50, p95, p99), not averages. Techniques: timeouts with sensible limits and a graceful fallback; streaming for perceived speed; shorter prompts and outputs; parallel independent steps (retrieval and moderation); and hedged requests: if a call has not returned after a threshold (say near your p90), send a second copy and use whichever answers first, cancelling the other. Hedging cuts the tail at the price of some extra requests (and cost), so apply it only to idempotent, cost-tolerant calls. The simulation shows the effect with invented numbers.
Hedging cuts the tail (simulation), run
I ran this with plain Python 3 (standard library only). This is a simulation with invented numbers and seeded random draws, so it shows the shape of the effect, not measurements of any product. The median stays 1.02 s, but p99 falls from 9.16 s to 2.95 s and p95 from 2.73 s to 2.25 s when a second copy is sent after 1.5 s. About 17% extra requests are sent.
import random
rng = random.Random(5)
def latency(): # mostly fast, with an occasional slow outlier (seconds)
return rng.lognormvariate(0, 0.35) if rng.random() > 0.05 else rng.uniform(6, 10)
def pct(values, p): v = sorted(values); return v[int(p * (len(v) - 1))]
N = 20000
single = [latency() for _ in range(N)]
hedged = [min(latency(), 1.5 + latency()) for _ in range(N)] # send a second copy after 1.5 s; take the first answer
for name, xs in (("single request", single), ("hedged after 1.5 s", hedged)):
print(f"{name:20} p50 {pct(xs, .5):.2f}s p95 {pct(xs, .95):.2f}s p99 {pct(xs, .99):.2f}s")
extra = sum(1 for x in single if x > 1.5) / N
print(f"extra requests sent by hedging: about {extra:.1%}")
Output:
single request p50 1.02s p95 2.73s p99 9.16s hedged after 1.5 s p50 1.02s p95 2.25s p99 2.95s extra requests sent by hedging: about 16.7%
Hedge only idempotent calls
Sending a duplicate of a side-effecting action could repeat its effect.
Quick check: Why track p95 and p99 rather than only the average?
- Averages are always wrong
- Percentiles are cheaper to compute
- Averages hide the slow requests that users notice
- Users ignore speed
Answer
Averages hide the slow requests that users notice — The tail determines perceived reliability.