Lesson 21 / 26

Running It: Availability, Config and Upgrades

Keep the gateway and mesh highly available and change them safely.

The front door must not fail

A gateway is a single point of failure by position, so run several instances across availability zones behind a load balancer, with health checks and autoscaling, and size connection limits and timeouts deliberately. Treat gateway and mesh configuration as code: keep it in Git, review changes, validate before applying (nginx -t for syntax; mesh analysers and dry-run for resources), and roll out in stages, because one bad rule can take down every API at once. Upgrade control planes and proxies gradually and keep version skew within the supported range. Document who owns the platform, how to roll back, and how to bypass the gateway or disable a policy in an emergency. Keep gateway logic thin so deployments stay rare and safe.

Validating config before applying, run

The gateway container started with nginx -t first; it reported the configuration ok before serving traffic. Build this check into CI and into startup.

nginx: the configuration file /etc/nginx/nginx.conf syntax is ok
nginx: configuration file /etc/nginx/nginx.conf test is successful

Practise the bypass

If the mesh control plane or gateway misbehaves, you need a tested way to keep serving traffic (for example, fail-open proxies or a direct path). Rehearse it before an incident.

Quick check: Why validate gateway configuration before applying it?

  • Validation makes traffic faster
  • One bad rule can break every API behind the gateway
  • Gateways cannot be configured otherwise
  • It replaces monitoring
Answer

One bad rule can break every API behind the gateway — The gateway is shared by many services, so mistakes have a wide blast radius.