Retry Policy
A service mesh configuration that automatically retries failed requests to a destination without application code changes. In Istio, retry policies are defined per route in a VirtualService, specifying the number of attempts, per-attempt timeout, and which error conditions trigger a retry.
What is a Retry Policy
Without a retry policy, if a network hiccup causes a request to fail, the calling service receives an error and must handle it in application code. With a retry policy, the sidecar proxy automatically retries the request transparently - the application may never know the first attempt failed.
Without retry policy:orders-service → request to payments-service fails (connection reset)orders-service receives error → must handle it in application code With retry policy (configured in VirtualService):orders-service → request to payments-service failsorders-service sidecar automatically retries 2 more timesthird attempt succeeds → response returned to orders-serviceorders-service never saw the failuresConfiguring Retries in Istio
Retry configuration lives inside the VirtualService, at the route level:
apiVersion: networking.istio.io/v1kind: VirtualServicemetadata: name: payments-service namespace: productionspec: hosts: - payments-service http: - route: - destination: host: payments-service subset: v1 timeout: 10s ## total time budget for all attempts combined retries: attempts: 3 ## try up to 3 times total (1 initial + 2 retries) perTryTimeout: 3s ## each individual attempt gets 3 seconds retryOn: 5xx,connect-failure,retriable-4xxWhat retryOn Conditions Mean
5xx: retry on any 5xx response codeconnect-failure: retry if connection to service failsretriable-4xx: retry on 409 Conflict responsesreset: retry on TCP resetgateway-error: retry on 502, 503, 504 responsesretriable-status-codes: specify exact status codesNot all errors should be retried:
- 4xx client errors (400, 401, 403, 404) should generally NOT be retried
- POST requests that are not idempotent should NOT be retried (could cause duplicate processing)
- 5xx server errors on idempotent endpoints can safely be retried
Retry Budget and Retries in Linkerd
Linkerd uses "retry budgets" instead of per-request retry counts. A retry budget allows a total percentage of additional requests to be retried:
## Linkerd retry budget (in ServiceProfile)spec: retryBudget: retryRatio: 0.2 ## allow up to 20% additional retries minRetriesPerSecond: 10 ttl: 10sThis prevents retry storms - if every request retries 3 times during an outage, your traffic amplifies 3x at exactly the worst moment. A budget caps the total retry volume.
Retry and Idempotency
Only retry idempotent operations safely:
Safe to retry: GET, HEAD, OPTIONS, PUT (replace), DELETEUnsafe to retry: POST, PATCH (without idempotency key) At Razorpay: payment initiation POST requests must NOT be retriedwithout an idempotency key - duplicate payment requests are unacceptable.GET requests for payment status are safe to retry.RememberThe
timeoutin a VirtualService is the TOTAL time budget for ALL attempts including retries. Withtimeout: 10sandattempts: 3withperTryTimeout: 3s- three successful 3-second attempts would take 9 seconds total, within the 10-second budget. A slowdown causing each attempt to hit 3.5 seconds would exhaust the budget at 2 attempts.
Common MistakeSetting retries on POST endpoints that create resources without checking for idempotency. If the first attempt succeeds but the response is lost in transit, a retry creates a duplicate resource. Use idempotency keys or only retry explicitly safe HTTP methods.
Frequently Asked Questions
How is a service mesh retry policy different from retry logic built into application code?
A mesh-level retry policy runs in the sidecar proxy (e.g., Envoy in Istio) and applies uniformly to every service without touching application code, whereas hand-rolled retries require every team to implement and maintain their own logic consistently — which in practice they rarely do. The mesh also has visibility other layers lack: it can retry based on actual network-level conditions (connection failures, specific HTTP status codes, gRPC status) defined declaratively per route.
Why can automatic retries make an outage worse instead of better?
Retrying failed requests to an already-overloaded or failing service multiplies the load hitting it — a retry storm — which can turn a partial degradation into a full outage as retries compound faster than the service can recover. Production retry policies should always pair with sane per-attempt timeouts, a capped retry budget, and ideally circuit breaking, rather than retrying indefinitely on every failure type.