When a Kubernetes cluster runs 40 microservices, getting packets between pods is the easy part. Kubernetes handles that. The hard part is making those connections secure, reliable, and observable — without writing the same retry logic, mTLS setup, and metrics collection code in every service, in every language. At Swiggy, a slow restaurant service should not bring down the order service. At Zerodha, traffic between internal financial services must be encrypted even inside the cluster. At Razorpay, releasing a new payment service version to 5% of traffic before committing to 100% is a hard business requirement. A service mesh solves all of this at the infrastructure layer, not the application layer.
What This Pillar Covers
- The cascading failure problem that raw Kubernetes networking cannot prevent
- The sidecar proxy pattern — how a second container per pod handles all networking transparently
- The control plane vs data plane distinction — what istiod does vs what Envoy does
- Mutual TLS between every service, automatic certificate rotation, with zero application code changes
- Canary releases and precise traffic splitting using VirtualService and DestinationRule
- Circuit breaking and outlier detection — ejecting failing pods before they cascade
- Observability through Kiali, Prometheus metrics, and distributed tracing built into the mesh
- Istio vs Linkerd — when each is the right choice based on operational trade-offs
Who This Is For
Platform engineers and senior DevOps engineers who manage Kubernetes clusters at scale and need to add security, reliability, and observability to service-to-service communication. This pillar assumes you are already comfortable with Kubernetes Deployments, Services, and namespaces.
Why This Matters in Production
At PhonePe, every UPI transaction passes through multiple internal services. Without mTLS, that data travels as plain text on the internal Kubernetes network. Without circuit breaking, a slow downstream service can cascade into a full system outage. Without traffic splitting, every deployment is a high-stakes all-or-nothing rollout. A service mesh moves all of this from application code into infrastructure — and makes it consistent, auditable, and manageable across every service in the cluster regardless of the language they are written in.
Prerequisites
- Kubernetes Deployments, Services, and namespaces
- Understanding of pods and how containers run inside them
- Basic familiarity with TLS and HTTPS
- kubectl proficiency for cluster interaction