Lesson 19 / 32
Circuit Breakers & Retries
Failing fast and isolating trouble instead of cascading it.
The cascading failure problem
When a downstream service slows down, callers pile up waiting threads/connections, which slows them down too — the failure cascades upstream and can take down the whole system.
A circuit breaker in your house
Like an electrical breaker, a circuit breaker pattern trips open after repeated failures — it stops calling the failing service immediately (fast failure instead of a hanging timeout), waits, then tries a trickle of test calls (half-open) before closing again.
Retry with backoff
Blind immediate retries make an overloaded service worse. Back off exponentially and add jitter so retries don't all land at once.
delay = base * 2^attempt + random_jitter
if attempt > max_attempts: give up, fallback
else: sleep(delay); retry
Output:
Prevents retry storms from making an outage worse
Bulkheads and graceful degradation
Bulkheads give each dependency its own thread pool/connection limit so one slow dependency can't starve requests to others — like a ship's watertight compartments. When a non-critical dependency is down, degrade gracefully: show cached recommendations instead of failing the whole page.