Lesson 22 / 26

Common Pitfalls and Failure Drills

Recognise typical mistakes: logic in the gateway, retry storms, silent policy gaps and sidecar surprises.

What usually goes wrong

(1) Business logic in the gateway: it becomes a hard-to-test monolith. (2) Retry storms: retries at several layers multiply load during an outage. (3) Timeouts missing or inverted: inner timeouts longer than outer ones. (4) Trusting client headers such as X-User-Id. (5) Policy gaps: a route that skips authentication because a rule was added in the wrong order. (6) Sidecar surprises: apps that start before their proxy is ready, long-lived connections that ignore rebalancing, or protocol detection issues. (7) Unbounded cardinality in metrics labels. (8) No bypass or rollback plan. Practise failure drills in staging: kill an upstream, slow one down, break a certificate, exhaust a pool, and check that alerts fire and behaviour matches what you expect.

A failure drill checklist

The demo environment is perfect for these drills: stop gwg-orders-v1, watch the gateway return 502/504, then restore it.

Drill                         Expect
stop one upstream             retries to a healthy peer; no client errors (if idempotent)
stop ALL upstreams            fast 502/503 (not a hang); circuit opens; alert fires
slow upstream (3 s)           gateway timeout (504) at the configured limit
expired / wrong certificate   mTLS handshake fails; calls rejected; alert on cert errors
burst of 100 requests         429 beyond rate+burst; backends stay healthy
bad config pushed             validation rejects it; or staged rollout catches it

Quick check: Which is a typical gateway anti-pattern?

  • Rate limiting
  • Terminating TLS there
  • Adding request IDs
  • Putting business logic in the gateway
Answer

Putting business logic in the gateway — Gateways should hold thin policy, not application rules.