Skip to main content

Mutual TLS (mTLS)

A two-way TLS authentication protocol where both the client and the server verify each other's identity before any communication begins. In a service mesh, mTLS is automatically applied between all services without application code changes.

What is Mutual TLS

Regular TLS is one-sided. When your browser connects to a bank website, you verify the bank's identity (via its certificate) but the bank does not cryptographically verify yours. Mutual TLS adds the other direction - both sides verify each other.

◈ DIAGRAM
Regular TLS:
Client connects → verifies Server certificate → encrypted tunnel
Server does NOT verify Client identity
Mutual TLS:
Client connects → verifies Server certificate
Server verifies Client certificate → encrypted tunnel
Both sides authenticated before any data flows

Why mTLS Matters Inside Kubernetes

Without mTLS, traffic between pods inside a Kubernetes cluster is plain text. Anyone with access to the cluster network - a compromised pod, a misconfigured network policy, a malicious container - can read that traffic.

◈ DIAGRAM
Without mTLS:
orders-service → (plain HTTP, readable by anyone) → payments-service
With mTLS:
orders-service → (encrypted, both sides verified) → payments-service

At Zerodha, internal services pass financial transaction data between them. Without mTLS, that data travels in plain text on the internal Kubernetes network - a serious compliance and security risk.

How mTLS Works Without Touching Your Code

The service mesh handles mTLS in the sidecar proxies. Your application code opens a plain HTTP connection to the target service. The sidecar intercepts that connection and performs the mTLS handshake with the target service's sidecar - transparently.

TEXT
Your app: opens plain HTTP to payments-service:8080
|
Your sidecar: intercepts, holds a cert from mesh CA
|
mTLS handshake: your sidecar verifies payments-service sidecar cert
| payments-service sidecar verifies your cert
|
Encrypted tunnel established between sidecars
|
payments-service sidecar: decrypts, delivers plain HTTP to payments app

Neither application knows about the mTLS handshake. Neither application code changes.

Configuring mTLS in Istio

YAML
## Enforce mTLS across the entire production namespace
apiVersion: security.istio.io/v1beta1
kind: PeerAuthentication
metadata:
name: default
namespace: production
spec:
mtls:
mode: STRICT ## only encrypted traffic accepted

Two modes:

TEXT
STRICT: only mTLS traffic accepted, plain HTTP rejected
PERMISSIVE: both mTLS and plain HTTP accepted

Use PERMISSIVE during migration (some services may not have sidecars yet). Switch to STRICT once all services have sidecars injected.

Certificate Management

Istio manages its own Certificate Authority (CA) built into istiod. Certificates are:

  • Issued automatically when a pod starts
  • Rotated every 24 hours automatically
  • Tied to the Kubernetes Service Account identity of the pod

You never manage these certificates manually.

Remember

mTLS in a service mesh encrypts the connection between sidecars - not between applications. Your application still opens plain HTTP. The encryption tunnel is between the two sidecar proxies. This is why no code changes are needed.

Security

Never enable STRICT mTLS on an existing cluster without first verifying all pods in the namespace have sidecars injected. Any pod without a sidecar will immediately lose the ability to communicate with other services in the namespace.

Frequently Asked Questions

How does mTLS in a service mesh differ from regular TLS on a public website?

Regular TLS (as used by HTTPS websites) only verifies the server's identity to the client — the browser checks the server's certificate, but the server has no cryptographic proof of who's connecting. mTLS requires both sides to present and validate certificates, so a service only accepts traffic from another service it can cryptographically verify. In meshes like Istio or Linkerd, a sidecar proxy handles certificate issuance, rotation, and the handshake transparently, so application code never touches TLS directly.

What's a common mistake teams make when rolling out mTLS across a cluster?

Switching mesh-wide mTLS enforcement straight to STRICT mode without first running in PERMISSIVE mode (which accepts both plaintext and mTLS) is a frequent cause of outages — any service or job still communicating without a sidecar (a legacy pod, a health check from outside the mesh) suddenly gets rejected. The safer rollout is PERMISSIVE first, verify traffic via mesh telemetry that everything's actually using mTLS, then flip to STRICT namespace by namespace.