Understanding the Problem Service Mesh Solves
You have 40 microservices running on Kubernetes. Payments, orders, inventory, notifications, user profiles - all talking to each other over the network. Everything works fine in staging with 5 services. In production with 40, things get messy fast.
The order service calls the payments service. The payments service is slow - maybe a database is struggling, maybe there is a bug in a new deployment. The order service waits. Then more requests pile up. The order service runs out of threads waiting for payments to respond. Now the order service is also slow. Every service that calls orders starts waiting too. Within minutes, a problem in one service has taken down the entire system. This is called a cascading failure.
Kubernetes gives you a way to run services. It does not give you any of the following:
- Automatic retries when a service call fails
- Encrypted communication between services inside the cluster
- The ability to send 5% of traffic to a new version before rolling it out fully
- A dashboard showing which services are slow and why
- Automatic circuit breaking that stops calls to a failing service
You could add all of this inside every service yourself. But that means writing the same networking logic in every service, in every programming language your teams use. When you need to change that logic, you change 40 services. This is where a service mesh steps in.
What raw Kubernetes networking gives you
Kubernetes networking handles one thing well - getting packets from one pod to another. Every pod gets an IP address. Services get a stable DNS name. Traffic routes from a Service to its matching pods through kube-proxy.
That is where Kubernetes stops. It does not know if the response was an error. It does not retry. It does not encrypt. It does not measure latency per service. It treats all traffic as equal regardless of which version a pod is running.
For two or three services this is fine. For 20 or 40, it is not enough.
Where service mesh fits in the platform engineering stack
If you have already set up Kubernetes deployments, you are adding a service mesh as a layer on top - it does not replace anything. Your pods keep running. Your services keep working. The mesh adds capabilities underneath your application code without requiring any changes to it.
Your application code |Service Mesh (Istio or Linkerd) |Kubernetes networking (kube-proxy, CNI) |Physical or cloud networkEverything your app does still works. The mesh intercepts traffic, applies policies, and collects data at the layer below your code.
Understanding the Sidecar Pattern
The key mechanism behind every service mesh is the sidecar proxy. Understanding this one pattern makes everything else click.
Think about a secretary who sits outside every meeting room. Every letter that goes into the room passes through the secretary. Every letter that comes out passes through the secretary too. The people inside the room do not know the secretary is there - they just send and receive letters as normal. But the secretary logs everything, checks credentials, can hold back letters if the room is too busy, and can route urgent letters to a backup room if the main one is unavailable.
That is exactly what a sidecar proxy does for your pods.
How the sidecar gets attached to every pod
When you install a service mesh on a cluster, it modifies how pods are created. It watches for new pods and automatically injects a second container into each one. This second container is the sidecar proxy - Envoy in Istio's case, a lightweight Rust-based proxy in Linkerd's case.
Pod (before mesh) Pod (after mesh)+------------------+ +------------------+| Your app | | Your app || container | | container |+------------------+ +------------------+ | Sidecar proxy | | (auto-injected)| +------------------+The sidecar and your app container share the same network namespace. This means the sidecar can intercept all incoming and outgoing traffic using iptables rules - without your app knowing. Your app thinks it is talking directly to the payments service. In reality, it is talking to its own sidecar, which talks to the payments service's sidecar, which then delivers the traffic to the payments app.
The control plane and data plane
A service mesh has two parts. Once you know what they do, reading Istio and Linkerd documentation becomes much clearer.
The data plane is all the sidecars running alongside your pods. This is where actual traffic flows. Sidecars intercept, forward, encrypt, and collect metrics on every request.
The control plane is the central brain. It watches the cluster for service changes, calculates routing rules, and pushes configuration down to all the sidecars. In Istio this is called istiod. In Linkerd this is called the control plane. Sidecars connect to the control plane at startup and receive their configuration from it.
The control plane never sees your actual traffic. It only pushes config. All traffic handling happens in the data plane sidecars.
Implementing mTLS Between Services
mTLS stands for mutual TLS. Regular TLS is one-sided - your browser verifies the server's identity, but the server does not verify yours. Mutual TLS means both sides verify each other's identity. Every service proves who it is before communication begins.
Without mTLS inside your cluster, traffic between pods is plain text. Anyone with access to the cluster network - a compromised pod, a malicious container, a misconfigured network policy - can read that traffic. At Zerodha, where internal services pass financial transaction data between them, unencrypted internal traffic is a serious compliance and security risk.
How mTLS works inside the mesh without touching your code
This is the part that seems like magic until you understand the sidecar pattern. Your application code opens a plain HTTP connection to the payments service. Your sidecar intercepts that connection. The sidecar holds a certificate issued by the mesh's certificate authority. The payments service's sidecar also holds a certificate from the same CA.
The two sidecars perform a TLS handshake between themselves, verifying each other's identities using those certificates. The encrypted tunnel is between the sidecars. Your app code never changes. It still opens plain HTTP. The encryption happens transparently below it.
## Enable mTLS for the entire production namespaceapiVersion: security.istio.io/v1beta1kind: PeerAuthenticationmetadata: name: default namespace: productionspec: mtls: mode: STRICT ## only encrypted traffic allowed, plain rejectedNote
STRICTmode means any service without a sidecar cannot communicate in this namespace. UsePERMISSIVEfirst when migrating an existing cluster so services without sidecars still work while you roll out injection gradually.
RemembermTLS certificates in Istio rotate automatically every 24 hours. You do not manage them manually. The control plane handles issuance and rotation.
Verifying mTLS is active
## Check that mTLS is enforced between two servicesistioctl x authz check \ $(kubectl get pod -l app=orders -n production \ -o jsonpath='{.items[0].metadata.name}') \ -n productionACTION AuthorizationPolicy RULESALLOW - -NoteIf mTLS is working, traffic between services will show as encrypted in the Kiali dashboard (lock icon on service graph edges).
Implementing Canary Releases with Traffic Splitting
A canary release is a deployment strategy where you send a small percentage of real traffic to a new version before fully rolling it out. The name comes from the old mining practice of sending a canary into a tunnel first - if it survives, it is safe for people.
At Swiggy, releasing a new version of the restaurant service to 100% of users at once is risky. If there is a bug, every user is affected. With a canary, you send 5% of traffic to v2, watch error rates and latency for 30 minutes, and only promote to 100% if metrics look healthy.
Without a service mesh
Without a mesh, doing traffic splitting requires running two separate Kubernetes Deployments and adjusting replica counts to approximate percentages. Want 5% to v2? Run 1 pod of v2 and 19 pods of v1. This is imprecise, expensive, and impossible to control below the granularity of your pod count.
With Istio VirtualService
With Istio, you define traffic weights explicitly and precisely. It does not matter how many pods each version has.
## Route 95% of traffic to v1, 5% to v2 of the restaurant serviceapiVersion: networking.istio.io/v1kind: VirtualServicemetadata: name: restaurant-service namespace: productionspec: hosts: - restaurant-service ## the Kubernetes Service name http: - route: - destination: host: restaurant-service subset: v1 ## defined in DestinationRule below weight: 95 - destination: host: restaurant-service subset: v2 weight: 5## Define which pods belong to v1 and which to v2apiVersion: networking.istio.io/v1kind: DestinationRulemetadata: name: restaurant-service namespace: productionspec: host: restaurant-service subsets: - name: v1 labels: version: "v1" ## pods with this label get v1 traffic - name: v2 labels: version: "v2" ## pods with this label get v2 trafficNote
VirtualServicecontrols how traffic is routed.DestinationRulecontrols which pods belong to each subset. You need both for traffic splitting to work.
To promote v2 after validation, change the weights to weight: 0 and weight: 100.
To roll back, delete the VirtualService and all traffic returns to the default Service routing.
Header-based routing for internal testing
## Send traffic to v2 only when request has header x-canary: true## For internal QA teams to test v2 without affecting real users http: - match: - headers: x-canary: exact: "true" route: - destination: host: restaurant-service subset: v2 - route: - destination: host: restaurant-service subset: v1Implementing Circuit Breaking
Circuit breaking is named after electrical circuit breakers. When too much current flows through a circuit, the breaker trips and stops the flow before the wiring catches fire. In software, a circuit breaker stops sending requests to a service that is failing - before those failed requests pile up and take down the calling service too.
Without circuit breaking, if the payments service is down and orders keeps retrying, orders fills up its thread pool waiting for responses. Orders becomes slow. Notification service, which calls orders, also fills up. Within minutes, a payments failure has cascaded into a full system outage.
With circuit breaking, after a configured number of failures, the mesh stops forwarding requests to payments entirely and returns an error immediately. The calling service can handle that fast error gracefully instead of hanging. When payments recovers, the circuit closes and traffic resumes.
Configuring circuit breaking in Istio
## Trip the circuit breaker after 5 consecutive errors## or if connection pool is exhaustedapiVersion: networking.istio.io/v1kind: DestinationRulemetadata: name: payments-circuit-breaker namespace: productionspec: host: payments-service trafficPolicy: connectionPool: http: http1MaxPendingRequests: 100 ## max queued before rejection maxRequestsPerConnection: 10 ## max per connection outlierDetection: consecutive5xxErrors: 5 ## trip after 5 consecutive 5xx interval: 10s ## check error rate every 10s baseEjectionTime: 30s ## eject failing pod for 30s maxEjectionPercent: 50 ## max 50% of pods ejectableRememberCircuit breaking in Istio works at the pod level, not the service level. If one pod in the payments Deployment is failing, that specific pod gets ejected. Healthy pods in the same Deployment keep receiving traffic. This is called outlier detection.
Common MistakeSetting
maxEjectionPercent: 100means Istio can eject all pods if they all start returning errors. This makes your service completely unreachable. Keep it at 50% or lower so at least half your pods always remain in rotation.
Observability with Kiali and Metrics
The third major benefit of a service mesh is observability you get for free. Because every request passes through a sidecar, the mesh can measure latency, error rates, and request volume for every service-to-service call - without any instrumentation in your application code.
Kiali is the standard observability dashboard for Istio. It shows a live graph of your services with traffic flowing between them. Each edge in the graph shows request rate, error rate, and latency. Each node shows health status.
What you can see in Kiali
- Which services are calling which other services (and which you did not know about)
- Where in the call chain a slow response is originating
- Which version of a service is receiving traffic during a canary rollout
- Which circuit breakers have tripped
- Whether mTLS is active on each connection (shown as a lock icon)
Metrics automatically collected by the sidecar
## Istio sidecars expose Prometheus metrics at port 15090## These are scraped automatically if you have Prometheus installed ## Example: check raw metrics from an orders pod sidecarkubectl exec -n production \ $(kubectl get pod -l app=orders -n production \ -o jsonpath='{.items[0].metadata.name}') \ -c istio-proxy -- \ curl -s localhost:15090/metrics | grep istio_requestsKey metrics the mesh provides out of the box:
| Metric | What it measures |
|---|---|
istio_requests_total |
Total request count, labelled by source, destination, status code |
istio_request_duration_milliseconds |
Latency histogram per service pair |
istio_tcp_connections_opened_total |
TCP connection rate for non-HTTP traffic |
TipInstall the Istio addons bundle to get Kiali, Prometheus, Grafana, and Jaeger all preconfigured together. One command:
kubectl apply -f https://raw.githubusercontent.com/istio/istio/release-1.22/samples/addons/kiali.yaml
Choosing Between Istio and Linkerd
The two most widely used service meshes are Istio and Linkerd. They solve the same core problems but make very different trade-offs. Choosing the wrong one adds unnecessary operational complexity.
| Dimension | Istio | Linkerd |
|---|---|---|
| Complexity | High - many CRDs, many config options | Low - focused feature set, fewer concepts |
| Resource overhead | Higher - Envoy proxy uses more CPU and RAM | Lower - Rust proxy is very lightweight |
| mTLS | Yes, automatic | Yes, automatic |
| Traffic splitting | Yes, via VirtualService | Yes, via HTTPRoute |
| Circuit breaking | Yes, via DestinationRule | Partial - retry budgets, no outlier detection |
| Observability | Kiali, Jaeger, Prometheus built in | Viz dashboard built in, lightweight |
| Multi-cluster | Yes, strong support | Yes, good support |
| Learning curve | Steep - Envoy concepts required | Gentle - simpler mental model |
| Best for | Large teams, complex routing needs, compliance requirements | Teams wanting fast setup, lower overhead, simpler ops |
| Protocol support | HTTP/1, HTTP/2, gRPC, TCP | HTTP/1, HTTP/2, gRPC, TCP |
When to choose Istio
Pick Istio when your team needs advanced traffic management - fault injection for chaos engineering, complex header-based routing, WebAssembly extensions for custom proxy behaviour, or strong multi-cluster federation. Istio is also the default choice when teams are already investing in the CNCF ecosystem and need the widest community support.
The trade-off is operational complexity. Istio has over 50 Custom Resource Definitions. A new engineer joining the team will need time to understand VirtualServices, DestinationRules, Gateways, PeerAuthentication, AuthorizationPolicies, and ServiceEntries before they can safely make changes.
When to choose Linkerd
Pick Linkerd when you want the core mesh capabilities - mTLS, retries, observability, traffic splitting - with the fastest installation path and lowest operational overhead. Linkerd's data plane proxy is written in Rust and uses significantly less CPU and memory than Envoy. For teams running on smaller clusters or with strict resource budgets, this matters.
Linkerd does not support some of the more advanced Istio features like fault injection or WebAssembly extensions. If your requirements are: secure service-to-service communication, automatic retries, and a service topology dashboard, Linkerd delivers all of this with far less configuration surface area.
RememberBoth meshes are CNCF projects. Linkerd is a graduated CNCF project. Istio became a CNCF graduated project in 2023. Either is a safe, production-grade choice. The decision is about operational trade-offs, not maturity.
Hands-On - Installing Istio and Running a Canary Release
This walkthrough installs Istio on a local Kubernetes cluster, deploys two versions of a service, and performs a live canary release by shifting traffic between versions.
Prerequisites: kubectl configured against a running cluster (Minikube, Kind, or cloud), and istioctl downloaded from istio.io/downloadIstio.
- Download and install the
istioctlCLI:
## Download Istio 1.22 or newer (the v1 networking APIs used in this lab need Istio 1.22+)curl -L https://istio.io/downloadIstio | ISTIO_VERSION=1.22.0 sh - ## Add istioctl to your PATH for this sessionexport PATH=$PWD/istio-1.22.0/bin:$PATH ## Verify the CLI is availableistioctl version- Install Istio onto your cluster using the demo profile:
## demo profile includes addons and is suitable for learning## production profile is leaner - use it for real clustersistioctl install --set profile=demo -y ## Verify all Istio components are runningkubectl get pods -n istio-systemNAME READY STATUS RESTARTSistio-ingressgateway-xxx 1/1 Running 0istiod-xxx 1/1 Running 0- Enable automatic sidecar injection for your namespace:
## Label namespace so Istio injects sidecars automatically## Every new pod created in this namespace will get a sidecarkubectl create namespace productionkubectl label namespace production istio-injection=enabled ## Verify the label is appliedkubectl get namespace production --show-labels- Deploy two versions of an orders service:
## Save this as orders-v1.yamlcat <<EOF | kubectl apply -f -apiVersion: apps/v1kind: Deploymentmetadata: name: orders-v1 namespace: productionspec: replicas: 2 selector: matchLabels: app: orders version: v1 template: metadata: labels: app: orders version: v1 ## DestinationRule uses this to identify v1 pods spec: containers: - name: orders image: nginx:1.24 ## stand-in for a real orders service ports: - containerPort: 80EOF ## Save this as orders-v2.yamlcat <<EOF | kubectl apply -f -apiVersion: apps/v1kind: Deploymentmetadata: name: orders-v2 namespace: productionspec: replicas: 2 selector: matchLabels: app: orders version: v2 template: metadata: labels: app: orders version: v2 spec: containers: - name: orders image: nginx:1.25 ## v2 runs a newer nginx version ports: - containerPort: 80EOF ## Create Service selecting all orders pods (both versions)cat <<EOF | kubectl apply -f -apiVersion: v1kind: Servicemetadata: name: orders namespace: productionspec: selector: app: orders ## selects all orders pods regardless of version ports: - port: 80 targetPort: 80EOF- Verify sidecars were injected (pods should show 2/2 containers):
kubectl get pods -n production ## Expected output shows 2 containers per pod (app + sidecar)NAME READY STATUS RESTARTSorders-v1-xxx-xxx 2/2 Running 0orders-v2-xxx-xxx 2/2 Running 0Note
2/2means both containers in the pod are running. The second container is the Istio sidecar proxy. If you see1/1, sidecar injection did not work - check that the namespace has theistio-injection=enabledlabel.
- Apply DestinationRule to define v1 and v2 subsets:
cat <<EOF | kubectl apply -f -apiVersion: networking.istio.io/v1kind: DestinationRulemetadata: name: orders namespace: productionspec: host: orders subsets: - name: v1 labels: version: v1 - name: v2 labels: version: v2EOF- Apply VirtualService to start with 100% traffic on v1:
cat <<EOF | kubectl apply -f -apiVersion: networking.istio.io/v1kind: VirtualServicemetadata: name: orders namespace: productionspec: hosts: - orders http: - route: - destination: host: orders subset: v1 weight: 100 - destination: host: orders subset: v2 weight: 0EOF- Shift 10% of traffic to v2 to begin the canary:
## Edit the VirtualService to shift 10% traffic to v2kubectl patch virtualservice orders -n production \ --type=json \ -p='[ {"op":"replace","path":"/spec/http/0/route/0/weight","value":90}, {"op":"replace","path":"/spec/http/0/route/1/weight","value":10} ]' ## Verify the change was appliedkubectl get virtualservice orders -n production -o yaml | grep weight- Monitor metrics and promote or roll back:
## Install Kiali to visualise the canary traffic splitkubectl apply -f \ https://raw.githubusercontent.com/istio/istio/\ release-1.22/samples/addons/kiali.yaml ## Open the Kiali dashboard in your browseristioctl dashboard kialiWatch the Kiali graph for v2 error rates and latency. If v2 looks healthy after your monitoring window, promote by setting v2 to 100%. If v2 shows problems, roll back by setting v1 back to 100%.
## Promote v2 to full production traffickubectl patch virtualservice orders -n production \ --type=json \ -p='[ {"op":"replace","path":"/spec/http/0/route/0/weight","value":0}, {"op":"replace","path":"/spec/http/0/route/1/weight","value":100} ]'- Clean up after the lab:
## Remove all lab resourceskubectl delete namespace production ## Uninstall Istio from the clusteristioctl uninstall --purge -ykubectl delete namespace istio-systemQuick Reference
| Resource | What it controls |
|---|---|
PeerAuthentication |
Enforces mTLS mode (STRICT or PERMISSIVE) per namespace or workload |
VirtualService |
Controls routing rules - weights, header matching, retries, timeouts |
DestinationRule |
Defines subsets (v1, v2) and traffic policies (circuit breaking, connection pools) |
AuthorizationPolicy |
Controls which services are allowed to talk to which |
istioctl analyze |
Checks your Istio config for errors before applying |
istioctl proxy-status |
Shows sync status of all sidecars with the control plane |
istioctl dashboard kiali |
Opens the Kiali service graph in your browser |
Common mistakes engineers make when working with service meshes:
Starting with STRICT mTLS on an existing cluster causes immediate outages for any service that does not have a sidecar injected yet. The fix is to start with PERMISSIVE mode, roll out sidecar injection incrementally across namespaces, verify all services work, and only then switch to STRICT. Changing the mode takes one kubectl apply - the migration plan is what takes time.
Forgetting that a VirtualService requires a matching DestinationRule to use subsets. If you reference subset v2 in a VirtualService but have not defined it in a DestinationRule, Istio will return 503 errors for all traffic matching that route. Always apply the DestinationRule before or alongside the VirtualService, and run istioctl analyze to catch this before it hits production.
Applying sidecar injection to the kube-system or istio-system namespaces by mistake. These system namespaces must not have sidecar injection enabled. Always label only your application namespaces. If you accidentally inject system namespaces, core cluster components can break in unexpected ways.
Assuming circuit breaking in Istio works like application-level circuit breakers (like Hystrix or Resilience4j). Istio's outlier detection works at the pod level and ejects unhealthy pods from the load balancing pool. It does not maintain a circuit state per calling service. A pod ejected because orders was calling it will also be ejected for calls from notifications. This is usually what you want, but it surprises engineers who expect per-caller circuit state.
Not setting resource requests and limits for Envoy sidecars on small clusters. Each Envoy sidecar takes roughly 50-100MB of memory and some CPU. In a cluster with 100 pods, that is 5-10GB of memory just for sidecars. On resource-constrained clusters, set global.proxy.resources in your Istio Helm values to bound sidecar resource usage before you hit OOM evictions.
Interview Questions This Topic Maps To
Explain how a service mesh implements mTLS without requiring application code changes. Walk through what happens at the network level when one service calls another.
Your team wants to do a canary release of the payments service. How would you use Istio to route 5% of traffic to the new version while ensuring instant rollback capability?
What is the difference between outlier detection and a traditional circuit breaker? When would Istio's outlier detection fail to protect a calling service?
A team is running 200 pods across 40 services. They want to add a service mesh. What would you check before enabling mTLS in STRICT mode, and what is the risk of skipping that step?