Traffic Splitting
A service mesh capability that routes a specified percentage of requests to different versions of a service simultaneously. Used for canary releases, A/B testing, and blue-green deployments without changing replica counts.
What is Traffic Splitting
Traffic splitting sends a portion of live requests to a new version of a service before fully rolling it out. This lets you validate the new version with real traffic and real users before committing to a full deployment.
All traffic → v1 (safe, proven, production) During canary:95% of traffic → v15% of traffic → v2 (new version being tested) After validation:100% of traffic → v2Without a Service Mesh
Without a service mesh, traffic splitting in Kubernetes requires adjusting replica counts to approximate percentages:
Want 5% to v2:Deploy 1 pod of v2, 19 pods of v1Kubernetes Service distributes requests proportionally5% → v2, 95% → v1 (approximately)Problems:
- 1 pod of v2 for 5% means you run 20 total pods for what should be 10
- You cannot get below 5% without running more than 20 pods
- Cannot control by header, user segment, or any other attribute
With Istio Traffic Splitting
Istio uses a VirtualService to set exact traffic weights and a DestinationRule to define which pods belong to which version:
## VirtualService: controls the traffic splitapiVersion: networking.istio.io/v1kind: VirtualServicemetadata: name: restaurant-service namespace: productionspec: hosts: - restaurant-service http: - route: - destination: host: restaurant-service subset: v1 weight: 95 - destination: host: restaurant-service subset: v2 weight: 5## DestinationRule: defines which pods are v1 and v2apiVersion: networking.istio.io/v1kind: DestinationRulemetadata: name: restaurant-service namespace: productionspec: host: restaurant-service subsets: - name: v1 labels: version: "v1" - name: v2 labels: version: "v2"The weights are exact - 5% of requests go to v2 regardless of how many pods each version has.
Header-Based Traffic Splitting
For internal testing without affecting customers:
http: - match: - headers: x-canary: exact: "true" route: - destination: host: restaurant-service subset: v2 - route: - destination: host: restaurant-service subset: v1QA engineers add the x-canary: true header in their requests. All other requests go to v1.
Rollback
To roll back a canary deployment, set v1 weight to 100 and v2 to 0. Or delete the VirtualService entirely to return to default Kubernetes Service routing.
RememberBoth VirtualService and DestinationRule are required for traffic splitting. VirtualService defines the weights and routes. DestinationRule defines which pods are in each subset. Applying one without the other causes 503 errors.
TipDuring a canary, watch error rate and p99 latency for v2 in Kiali. If both look healthy after your monitoring window, gradually increase v2 weight: 5% → 25% → 50% → 100%. If errors spike at any step, immediately set v2 weight to 0.
Frequently Asked Questions
How does mesh-based traffic splitting differ from just scaling replica counts to shift load?
Adjusting replica counts (e.g. 9 v1 pods vs 1 v2 pod for a rough 90/10 split) only gives you a coarse, resource-coupled approximation tied to however many pods happen to be running. A service mesh (Istio, Linkerd) splits traffic at the request-routing layer independently of replica count — you can send exactly 5% of traffic to v2 with a single replica each, and adjust the percentage without touching your Deployment's scale at all.
What's a common pitfall when running a canary with traffic splitting?
Splitting traffic by request count without also correlating error rates and latency per version — a canary can silently 5x its error rate on 5% of traffic and go unnoticed if dashboards aggregate both versions together. Pair traffic splitting with per-version metrics and automated rollback (e.g. Flagger, Argo Rollouts) rather than manually watching percentages, since a human rarely reacts fast enough to a bad canary at 2am.