Lesson 20 / 26

Choosing a Gateway or Mesh

Compare the popular options on features, operations and cost.

No universal winner

For gateways: NGINX and HAProxy are fast, proven reverse proxies configured by files; Kong and Apache APISIX (built on NGINX/OpenResty) add plugins, an admin API and a developer ecosystem; Envoy is a programmable proxy powering many gateways and meshes; Traefik is popular for container and Kubernetes auto-discovery; managed gateways (Amazon API Gateway, Azure API Management, Google Cloud API Gateway/Apigee) remove operations at the price of vendor lock-in and per-request cost. For meshes: Istio is the most featureful (and complex), Linkerd prioritises simplicity and low overhead, Consul suits mixed VM and Kubernetes estates, and Cilium uses eBPF for networking and optional mesh features. Choose by team skills, platform (Kubernetes or not), compliance needs and how much operational work you can own. Features and product names change quickly, so check current documentation.

Tools, trade-offs, pitfalls

Pick the simplest tool that meets the need and plan how you will run, upgrade and debug it.

Three questions: which tool, how to run, what can go wrong.
Figure 6.1 — Which tool, how to run and what can go wrong.

A decision guide

A starting point to narrow the options; validate with a small proof of concept.

Need                                           Reasonable starting choice
Simple routing + TLS + rate limit, few services  NGINX / HAProxy / Traefik
Public API with keys, plans, portal, analytics   Kong / APISIX / managed API gateway
No ops team, all-in on one cloud                 managed cloud API gateway
Kubernetes, want mTLS + retries + canaries       Linkerd (simple) or Istio (feature-rich)
Mixed VMs + Kubernetes                           Consul
< ~10 services, one team                         probably no mesh at all

Quick check: When is a service mesh probably unnecessary?

  • With hundreds of services in many languages
  • With a handful of services owned by one team
  • When mTLS everywhere is mandated
  • When teams need uniform retries and telemetry
Answer

With a handful of services owned by one team — A mesh adds complexity that pays off at larger scale and with polyglot teams.