Service Mesh
A dedicated infrastructure layer that handles all service-to-service communication in a microservices architecture. It adds security, observability, and reliability to inter-service traffic without requiring any changes to application code.
Service Mesh — Extended Technical Detail
What Is a Service Mesh in Simple Terms?
When you run 20 or 40 microservices on Kubernetes, every service needs to talk to other services reliably and securely. Without a service mesh, each team implements their own retry logic, their own timeout handling, their own metrics collection, and their own encryption — in every language their service is written in. Change the retry policy and you update 40 services.
A service mesh moves all of that networking logic out of application code and into a dedicated infrastructure layer that runs alongside your services transparently.
What a Service Mesh Provides
- Security — mTLS between every service, with automatic certificate rotation
- Observability — latency, error rate, and request volume per service pair, without app-side instrumentation
- Reliability — automatic retries, timeouts, and circuit breaking applied uniformly
- Traffic control — weighted routing, canary releases, and header-based routing
None of these require any changes to the applications themselves. The mesh intercepts traffic at the network layer using the sidecar proxy pattern — a second container injected into every pod that transparently handles all inbound and outbound traffic.
How It Works — Data Plane and Control Plane
A service mesh has two parts that work together:
Data Plane — the sidecar proxies running alongside every pod. They handle actual traffic flow, enforce mTLS, collect metrics, and apply routing rules received from the control plane.
Control Plane — the central brain (istiod in Istio). It pushes configuration to all sidecars and manages certificate issuance and rotation. Critically, the control plane never sees actual service traffic — only the sidecars do.
When to Use a Service Mesh
A service mesh adds real operational complexity. It is the right choice when:
- You have 10 or more services and managing networking logic per-service is painful
- You need mTLS between services for security or compliance requirements
- You need traffic splitting for safe deployments, such as canary releases
- You need unified observability across all service-to-service communication
For two or three services, a service mesh is very likely over-engineering — the operational cost outweighs the benefit at that scale.
Istio vs Linkerd
The two most widely used service meshes on Kubernetes take different approaches to the same problem:
| Aspect | Istio | Linkerd |
|---|---|---|
| Sidecar proxy | Envoy (C++) | Rust-based, lightweight |
| Complexity | High — 50+ CRDs | Low — focused feature set |
| Resource overhead | Higher (~100MB/sidecar) | Lower (~10-20MB/sidecar) |
Istio is generally the choice for large teams needing advanced routing and fault injection. Linkerd fits teams that want fast setup and lower resource overhead.
Resource Overhead — Plan Before You Install
Each sidecar proxy consumes real memory and CPU on every node it runs on:
Envoy (Istio): ~50-100 MB memory, ~0.1-0.5 vCPU per sidecarLinkerd proxy: ~10-20 MB memory, ~0.05 vCPU per sidecar 100 pods with Istio: 5-10 GB additional cluster-wide memory100 pods with Linkerd: ~1-2 GB additional cluster-wide memoryAt Hotstar or Swiggy scale, installing a mesh across hundreds of pods without accounting for this overhead can silently eat a meaningful chunk of cluster capacity before any application workload runs.
Quick Reference & Troubleshooting Commands
| Problem | Cause | Fix |
|---|---|---|
| Service-to-service calls suddenly fail after mesh install | Sidecar not injected into pod | Check namespace has istio-injection=enabled label, then restart the pod |
| Unexpected latency after adopting mesh | Sidecar proxy overhead uncounted in capacity planning | Re-check node CPU/memory headroom for the added sidecar cost |
| mTLS errors between services | One side has no sidecar (mixed mesh/non-mesh traffic) | Use PERMISSIVE mTLS mode during migration, STRICT only once fully rolled out |
| Traffic not routing per canary weights | VirtualService applied without matching DestinationRule subsets | Ensure DestinationRule subsets exist before referencing them in VirtualService |
RememberA service mesh solves networking problems — it is not a replacement for Kubernetes Services, Ingress, or your application's own business logic. It sits below your application code and above the raw Kubernetes network layer.
TipStart with Linkerd if your primary goal is mTLS and basic observability with minimal operational overhead. Move to Istio only when you need advanced traffic management like fault injection, fine-grained circuit breaking, or complex multi-cluster routing.
Common MistakeInstalling a service mesh on a small cluster without accounting for sidecar memory. Each Envoy sidecar takes 50-100 MB. On a 100-pod cluster that is 5-10 GB of overhead just for the mesh layer — enough to force unplanned node scale-up before a single new application pod is deployed.
SecurityA service mesh is not a substitute for NetworkPolicies. mTLS encrypts and authenticates traffic between meshed services, but without NetworkPolicy enforcement at the CNI layer, any pod can still attempt to reach any other pod's exposed ports. Use both together for real isolation on multi-tenant clusters.
Frequently Asked Questions
What problem does a service mesh actually solve that a load balancer doesn't?
A load balancer handles north-south traffic (client to cluster); a service mesh handles east-west traffic — the hundreds of calls services make to each other internally. Once you have more than a handful of microservices, you need consistent mTLS, retries, timeouts, and circuit breaking between every pair of them, and doing that per-service in application code doesn't scale. Istio and Linkerd are the two most widely deployed meshes on Kubernetes.
When should a team NOT adopt a service mesh?
If you're running fewer than ~10-15 services, or your team doesn't already have solid Kubernetes operational maturity, a mesh often adds more complexity than it removes — sidecar proxies add latency, memory overhead per pod, and a genuinely difficult new failure mode to debug (mesh-level connection resets that look like application bugs). Many teams get 80% of the value from just using a good ingress controller plus application-level retries first.