# Failure Handling in Layers — LLM Engineering Foundations

Source: https://www.geekswithgeeks.com/en/llm-engineering/r-failure

> Combine retries, timeouts, fallbacks and graceful degradation.

## Plan each failure, then test it

Typical failures and a layered response: **transient provider errors and rate limits** → retry with exponential backoff and jitter, within a total time budget; **provider outage** → fail over to another model or provider (keep prompts portable and test the fallback regularly), or a cached or simpler answer; **slow responses** → timeouts, streaming, hedging for idempotent calls; **invalid output** → bounded repair loop, then safe default or human queue; **retrieval empty or weak** → say "I could not find this" rather than guess; **tool errors** → return clear error text to the model, limit retries, never repeat side-effecting actions blindly; **runaway loops** → step, token, time and cost caps; **overload** → queue, shed low-priority work, return a clear "busy" message. Add a **circuit breaker** to stop calling a failing dependency for a short time. Then run **failure drills** in staging: disable the provider, inject slow responses and bad outputs, and confirm users get a sensible experience.

## Failure modes and responses

Fill in your own system; test each row in staging.

```text
Failure                         Layered response
rate limit / 5xx / timeout       retry with backoff+jitter inside a time budget -> fallback model -> cached/simple answer
provider outage                 circuit breaker -> secondary provider (portable prompts) -> graceful message
invalid structured output       bounded repair (<= 2) -> safe default or human queue; count the failure
empty or weak retrieval         "I could not find this in our documents" + offer a human; do not guess
tool error                      clear error text to the model; capped retries; no blind repeat of side effects
runaway agent loop              step/token/time/cost caps -> stop, report what was tried
overload                        queue, shed low priority, "busy, try again" with retry hints
```

## Run a failure drill each quarter

Disable the provider in staging and confirm users get a sensible fallback.

**Quiz:** What is the safest response when retrieval returns nothing relevant?

- [ ] Invent a plausible answer
- [x] Say the answer was not found, and offer a human
- [ ] Return an empty page
- [ ] Crash the app

*Answer:* Say the answer was not found, and offer a human. An honest "not found" beats a confident guess.
