Skip to main content

Fault Injection

A service mesh feature that deliberately introduces delays or error responses into traffic to test how services respond to failures, without requiring any changes to the services themselves. In Istio, configured via a VirtualService fault block.

What is Fault Injection

Fault injection is controlled chaos engineering. Instead of waiting for real failures to discover how your system handles them, you intentionally inject failures in a controlled way to test your resilience mechanisms.

A service mesh makes this possible without touching application code. The sidecar proxy introduces the fault at the network level - the target service never receives the request (for abort faults) or receives it with an artificial delay (for delay faults).

Why Fault Injection Matters

Without fault injection, you discover problems when they happen in production:

◈ DIAGRAM
Real failure scenario (discovered at 2 AM):
Payments service becomes slow → orders service has no timeout →
orders fills up thread pool → orders goes down → alerts fire
With fault injection (discovered during testing):
Inject 5-second delay into payments service →
observe: does orders service timeout correctly?
does orders service circuit break?
does the user experience degrade gracefully?

Two Types of Faults

Delay Fault: Adds artificial latency before forwarding the request

YAML
apiVersion: networking.istio.io/v1
kind: VirtualService
metadata:
name: payments-service
namespace: production
spec:
hosts:
- payments-service
http:
- fault:
delay:
percentage:
value: 10.0 ## inject delay on 10% of requests
fixedDelay: 5s ## add 5 seconds of delay
route:
- destination:
host: payments-service
subset: v1

Abort Fault: Returns an HTTP error response without forwarding the request

YAML
- fault:
abort:
percentage:
value: 5.0 ## abort 5% of requests
httpStatus: 503 ## return 503 Service Unavailable
route:
- destination:
host: payments-service
subset: v1

Combining Delay and Abort

You can combine both in a single fault block:

YAML
- fault:
delay:
percentage:
value: 50.0
fixedDelay: 2s
abort:
percentage:
value: 10.0
httpStatus: 500

50% of requests get a 2-second delay. 10% of requests get aborted with 500. Use this to simulate a degraded, partially unavailable service.

Fault Injection for Testing Circuit Breakers

Fault injection and circuit breaking are designed to work together:

TEXT
Step 1: Configure circuit breaker in DestinationRule
consecutive5xxErrors: 5, baseEjectionTime: 30s
Step 2: Inject 503 errors via VirtualService fault
abort 100% of requests with httpStatus: 503
Step 3: Observe circuit breaker trips after 5 errors
Kiali shows the pod ejected from rotation
Step 4: Remove fault injection
observe: does the circuit close and traffic resume?

This validates your resilience configuration works before a real outage.

Fault Injection is Istio Only

Linkerd does not support fault injection natively. This is one of the reasons to choose Istio over Linkerd - when your team does chaos engineering or needs to test failure scenarios without modifying application code.

Remember

Fault injection in a VirtualService affects ALL callers of that service - not just your test requests. Injecting a delay on 100% of requests to payments-service will slow down every service in your cluster that calls it. Use low percentages (1-5%) for production testing, or test in a dedicated test namespace.

Security

Never leave fault injection configured in a production VirtualService after a test. The fault configuration looks identical to normal routing configuration and is easy to forget. Always remove or disable fault blocks after completing your chaos engineering test.

Frequently Asked Questions

What two types of faults can typically be injected via a service mesh's fault injection feature?

Istio's VirtualService fault block supports two types: delay injection, which adds artificial latency before a request reaches its destination, and abort injection, which returns a specific HTTP or gRPC error code immediately instead of forwarding the request. Both can target a percentage of traffic, letting you test how downstream services degrade under partial failure without touching any application code.

Why is fault injection typically done at the mesh layer instead of inside application test code?

Doing it at the proxy layer means the same fault-injection config works against real, unmodified service binaries in staging or even production canaries, revealing issues in actual retry logic, timeouts, and circuit breakers rather than a mocked-out test double. A common mistake is only testing faults in unit tests with mocks, which validates the code's intent but not how the real network client (with its real timeout and connection-pool settings) actually behaves under injected latency.