Circuit Breaking
A resilience pattern that stops sending requests to a failing service after a configured number of errors, giving the service time to recover while preventing cascading failures from spreading to other services. In Istio, implemented via outlier detection in a DestinationRule.
What is Circuit Breaking
The name comes from electrical circuit breakers. In an electrical system, too much current causes a breaker to trip and stop the flow before wiring catches fire. In software, too many failed requests to a service causes the circuit breaker to trip - stopping new requests before they pile up and take down the calling service too.
Without circuit breaking:Payments service is failingOrders service keeps sending requestsOrders fills up its thread pool waiting for timeoutsOrders becomes slowNotification service calls Orders - it also fills upEntire system degrades in a cascade With circuit breaking:Payments service is failingAfter 5 consecutive errors, circuit tripsOrders service gets immediate error instead of waitingOrders handles the fast error gracefullyNo cascade. System remains stable.How Istio Implements Circuit Breaking
Istio calls this "outlier detection" and implements it at the pod level inside a DestinationRule:
apiVersion: networking.istio.io/v1kind: DestinationRulemetadata: name: payments-circuit-breaker namespace: productionspec: host: payments-service trafficPolicy: connectionPool: http: http1MaxPendingRequests: 100 ## max queued before rejection maxRequestsPerConnection: 10 ## max per keep-alive connection outlierDetection: consecutive5xxErrors: 5 ## trip after 5 consecutive failures interval: 10s ## check every 10 seconds baseEjectionTime: 30s ## eject failing pod for 30 seconds maxEjectionPercent: 50 ## never eject more than 50% of podsPod-Level Ejection
Istio's circuit breaking works at the pod level - not the service level. When a pod returns too many errors, that specific pod is ejected from the load balancing pool. Other healthy pods in the same service continue receiving traffic.
payments-service has 3 pods: A, B, CPod C starts returning 5xx errorsAfter threshold: Pod C ejected from rotationPods A and B continue serving traffic normallyAfter baseEjectionTime: Pod C re-added to poolIf Pod C is still failing: ejected again for 2× the timemaxEjectionPercent Warning
Setting maxEjectionPercent: 100 means Istio can eject ALL pods if they all fail. This makes your service completely unreachable. Keep it at 50% or lower so at least half your pods always remain in rotation.
Istio vs Application-Level Circuit Breakers
| Istio Outlier Detection | Hystrix / Resilience4j | |
|---|---|---|
| Where it runs | Sidecar proxy | Inside application code |
| Scope | Pod level - affects all callers | Per-caller circuit state |
| Code changes | None | Requires library integration |
| Granularity | All callers see same ejected pod | Each caller has independent circuit |
RememberIstio's outlier detection ejects pods for all callers simultaneously. If pod C is ejected because orders service triggered it, notification service also stops routing to pod C automatically. This is different from Hystrix where each caller maintains its own circuit state.
TipSet baseEjectionTime to at least 30 seconds. If a pod is failing due to a memory issue or a stuck database connection, 5-10 seconds is not enough time for it to recover. 30 seconds gives the pod time to either fix itself or be replaced by a liveness probe restart.
Frequently Asked Questions
How is circuit breaking different from a plain retry policy?
A retry policy responds to failed requests by trying again, which can worsen things for an already-struggling service by piling on load exactly when it can least handle it. Circuit breaking instead tracks failure patterns over a window and, once a threshold is crossed, stops sending requests to that destination for a cooldown period — protecting the failing service from further load and callers from wasting time on requests likely to fail. The two are complementary: retries handle transient blips, circuit breakers handle sustained degradation.
What's a common misconfiguration with circuit breaking in Istio?
Setting outlier detection thresholds (`consecutiveErrors`, `interval`) without accounting for legitimate traffic spikes, causing a healthy-but-momentarily-slow service to get ejected from the load balancing pool unnecessarily — which then concentrates traffic on the remaining instances and can cascade into ejecting those too. Thresholds tuned only against synthetic load tests, without validating against real production traffic variance, are a common source of circuit breakers that trip during normal peak hours rather than genuine outages.