Skip to main content

Circuit Breaking

A resilience pattern that stops sending requests to a failing service after a configured number of errors, giving the service time to recover while preventing cascading failures from spreading to other services. In Istio, implemented via outlier detection in a DestinationRule.

What is Circuit Breaking

The name comes from electrical circuit breakers. In an electrical system, too much current causes a breaker to trip and stop the flow before wiring catches fire. In software, too many failed requests to a service causes the circuit breaker to trip - stopping new requests before they pile up and take down the calling service too.

TEXT
Without circuit breaking:
Payments service is failing
Orders service keeps sending requests
Orders fills up its thread pool waiting for timeouts
Orders becomes slow
Notification service calls Orders - it also fills up
Entire system degrades in a cascade
With circuit breaking:
Payments service is failing
After 5 consecutive errors, circuit trips
Orders service gets immediate error instead of waiting
Orders handles the fast error gracefully
No cascade. System remains stable.

How Istio Implements Circuit Breaking

Istio calls this "outlier detection" and implements it at the pod level inside a DestinationRule:

YAML
apiVersion: networking.istio.io/v1
kind: DestinationRule
metadata:
name: payments-circuit-breaker
namespace: production
spec:
host: payments-service
trafficPolicy:
connectionPool:
http:
http1MaxPendingRequests: 100 ## max queued before rejection
maxRequestsPerConnection: 10 ## max per keep-alive connection
outlierDetection:
consecutive5xxErrors: 5 ## trip after 5 consecutive failures
interval: 10s ## check every 10 seconds
baseEjectionTime: 30s ## eject failing pod for 30 seconds
maxEjectionPercent: 50 ## never eject more than 50% of pods

Pod-Level Ejection

Istio's circuit breaking works at the pod level - not the service level. When a pod returns too many errors, that specific pod is ejected from the load balancing pool. Other healthy pods in the same service continue receiving traffic.

TEXT
payments-service has 3 pods: A, B, C
Pod C starts returning 5xx errors
After threshold: Pod C ejected from rotation
Pods A and B continue serving traffic normally
After baseEjectionTime: Pod C re-added to pool
If Pod C is still failing: ejected again for 2× the time

maxEjectionPercent Warning

Setting maxEjectionPercent: 100 means Istio can eject ALL pods if they all fail. This makes your service completely unreachable. Keep it at 50% or lower so at least half your pods always remain in rotation.

Istio vs Application-Level Circuit Breakers

Istio Outlier Detection Hystrix / Resilience4j
Where it runs Sidecar proxy Inside application code
Scope Pod level - affects all callers Per-caller circuit state
Code changes None Requires library integration
Granularity All callers see same ejected pod Each caller has independent circuit
Remember

Istio's outlier detection ejects pods for all callers simultaneously. If pod C is ejected because orders service triggered it, notification service also stops routing to pod C automatically. This is different from Hystrix where each caller maintains its own circuit state.

Tip

Set baseEjectionTime to at least 30 seconds. If a pod is failing due to a memory issue or a stuck database connection, 5-10 seconds is not enough time for it to recover. 30 seconds gives the pod time to either fix itself or be replaced by a liveness probe restart.

Frequently Asked Questions

How is circuit breaking different from a plain retry policy?

A retry policy responds to failed requests by trying again, which can worsen things for an already-struggling service by piling on load exactly when it can least handle it. Circuit breaking instead tracks failure patterns over a window and, once a threshold is crossed, stops sending requests to that destination for a cooldown period — protecting the failing service from further load and callers from wasting time on requests likely to fail. The two are complementary: retries handle transient blips, circuit breakers handle sustained degradation.

What's a common misconfiguration with circuit breaking in Istio?

Setting outlier detection thresholds (`consecutiveErrors`, `interval`) without accounting for legitimate traffic spikes, causing a healthy-but-momentarily-slow service to get ejected from the load balancing pool unnecessarily — which then concentrates traffic on the remaining instances and can cascade into ejecting those too. Thresholds tuned only against synthetic load tests, without validating against real production traffic variance, are a common source of circuit breakers that trip during normal peak hours rather than genuine outages.