Skip to main content

kube-proxy

A network proxy that runs on every Kubernetes node, maintaining iptables or IPVS rules to enable Service-based load balancing and routing of network traffic to the correct pod endpoints across the cluster.

kube-proxy — Extended Technical Detail

What is kube-proxy in Simple Terms?

When you create a Kubernetes Service with a ClusterIP, that IP address is virtual — no actual network interface has it. kube-proxy is the component that makes that virtual IP work by programming routing rules on every node so traffic sent to the Service IP gets forwarded to a real, live pod IP.

◈ DIAGRAM
+------------------------------------------+
| Service created: payments-svc |
| ClusterIP: 10.96.45.200, port 3000 | <- Virtual IP — no interface owns this
+------------------------------------------+
|
v
+------------------------------------------+
| kube-proxy watches API server | <- Detects new Service and its Endpoints
+------------------------------------------+
|
v
+------------------------------------------+
| kube-proxy writes rules on EVERY node |
| |
| 10.96.45.200:3000 -> 10.244.1.15:3000 | <- pod-1 on mumbai-worker-1
| -> 10.244.2.8:3000 | <- pod-2 on mumbai-worker-2
| -> 10.244.3.22:3000 | <- pod-3 on mumbai-worker-3
+------------------------------------------+
|
v
+------------------------------------------+
| Any pod on any node reaches Service IP | <- Transparent load balancing
+------------------------------------------+

kube-proxy Modes — iptables vs IPVS

◈ DIAGRAM
+------------------------+ +------------------------------+
| iptables mode | | IPVS mode |
| (default) | | (high-performance) |
| | | |
| Linear chain scan | <------> | Hash table lookup |
| O(n) per packet | | O(1) per packet |
| | | |
| Sufficient up to | | Required for 1000+ Services |
| ~1000 Services | | (PhonePe, Razorpay scale) |
| | | |
| No load balance algos | | Round-robin, least-conn, |
| | | source-hash available |
+------------------------+ +------------------------------+

Check kube-proxy Status and Mode

Bash
# kube-proxy runs as a DaemonSet — one pod per node
kubectl get pods -n kube-system -l k8s-app=kube-proxy -o wide
# Output:
# NAME READY STATUS NODE
# kube-proxy-4xj9p 1/1 Running mumbai-worker-1
# kube-proxy-7kmnz 1/1 Running mumbai-worker-2
# kube-proxy-9plvw 1/1 Running mumbai-worker-3
# Check which mode kube-proxy is running in
kubectl logs -n kube-system kube-proxy-4xj9p | grep -i "using\|proxier"
# Output: Using iptables Proxier
# View the iptables rules kube-proxy has written for a Service
iptables -t nat -L KUBE-SERVICES | grep 10.96.45.200

Switching kube-proxy to IPVS Mode

YAML
# Edit the kube-proxy ConfigMap to switch to IPVS
kubectl edit configmap kube-proxy -n kube-system
# Find and update the mode field:
apiVersion: v1
kind: ConfigMap
metadata:
name: kube-proxy
namespace: kube-system
data:
config.conf: |
apiVersion: kubeproxy.config.k8s.io/v1alpha1
kind: KubeProxyConfiguration
mode: "ipvs" # Change from "" or "iptables" to "ipvs"
ipvs:
scheduler: "rr" # Round-robin (options: rr, lc, dh, sh, sed, nq)
Bash
# After editing the ConfigMap, restart all kube-proxy pods to pick up the change
kubectl rollout restart daemonset kube-proxy -n kube-system
# Verify the switch worked
kubectl logs -n kube-system kube-proxy-4xj9p | grep -i "using\|proxier"
# Output: Using ipvs Proxier
# View the IPVS virtual server table
ipvsadm -Ln | grep 10.96.45.200

How kube-proxy Handles Pod Churn

When pods are added or removed (deployments scale up/down, rolling updates), kube-proxy updates rules in real time:

Bash
+------------------------------------------+
| Deployment scales from 3 to 5 replicas | <- kubectl scale or HPA triggers
+------------------------------------------+
|
v
+------------------------------------------+
| API Server updates Endpoints object | <- New pod IPs added to endpoint slice
+------------------------------------------+
|
v
+------------------------------------------+
| kube-proxy detects Endpoints change | <- Watches API server continuously
+------------------------------------------+
|
v
+------------------------------------------+
| iptables/IPVS rules updated on all nodes | <- New pods immediately receive traffic
+------------------------------------------+

What kube-proxy Does NOT Handle

◈ DIAGRAM
+-----------------------------+
| kube-proxy handles | <- ClusterIP, NodePort, LoadBalancer (internal rules)
+-----------------------------+
| NOT handled by kube-proxy | <- Ingress traffic routing
+-----------------------------+
| NOT handled by kube-proxy | <- DNS resolution (that is CoreDNS)
+-----------------------------+
| NOT handled by kube-proxy | <- NetworkPolicy enforcement (that is CNI: Calico/Cilium)
+-----------------------------+
| NOT handled by kube-proxy | <- Pod-to-pod routing across nodes (that is CNI overlay)
+-----------------------------+

Troubleshooting Common kube-proxy Problems

Problem Symptom Fix
Service ClusterIP unreachable curl 10.96.45.200:3000 times out from inside a pod kube-proxy pod may be crashing — check kubectl logs -n kube-system kube-proxy-xxxxx
kube-proxy pod in CrashLoopBackOff Logs show failed to sync iptables rules Check if the node has iptables or ip_tables kernel module loaded: `lsmod
IPVS mode not activating Logs show fallback to iptables despite config change IPVS kernel modules not loaded — run modprobe ip_vs ip_vs_rr ip_vs_wrr ip_vs_sh on the node
Stale routing after pod deletion Traffic to deleted pod IP causes connection refused kube-proxy sync delay — check --iptables-sync-period flag (default 30s); reduce to 10s for high-churn workloads
NodePort not accessible externally Internal ClusterIP works but NodePort times out Firewall or security group blocking the NodePort range (30000-32767) — open the range on the cloud provider
Tip

For high-traffic clusters like PhonePe's payment processing nodes handling 1000+ Services, switch kube-proxy to IPVS mode. IPVS uses hash tables instead of linear iptables chains — at 5000 Services, iptables packet processing can add milliseconds of latency per request while IPVS stays at sub-millisecond.

Remember

kube-proxy does not handle Ingress traffic. It only manages internal cluster Service routing (ClusterIP, NodePort, LoadBalancer backend rules). External traffic routing from the internet into the cluster is handled by the Ingress Controller (NGINX, Traefik) and the cloud provider's load balancer.

Security

On Razorpay's production cluster, kube-proxy's iptables rules are a critical security boundary. Any node compromise that gives an attacker root access allows them to modify iptables rules directly — redirecting Service traffic to attacker-controlled pods. Use NodeRestriction admission and audit logging on the kube-proxy ConfigMap to detect tampering.

Common Mistake

Disabling kube-proxy entirely when adopting eBPF-based CNIs (like Cilium) without enabling Cilium's kube-proxy replacement mode first. Removing kube-proxy with no replacement leaves all Service ClusterIPs non-functional — a full cluster networking outage. Always enable kubeProxyReplacement: strict in Cilium config before removing the kube-proxy DaemonSet.

Frequently Asked Questions

What's the actual difference between kube-proxy's iptables mode and IPVS mode?

iptables mode implements Service load balancing as a chain of sequential rule matches — as Service and endpoint counts grow into the thousands, packet processing time grows roughly linearly since the kernel walks the rule chain. IPVS mode uses a hash table lookup instead, giving near-constant-time routing regardless of Service count, plus more load-balancing algorithms (round robin, least connection) beyond iptables' effectively random selection. IPVS is the recommended mode for large clusters.

What's a common source of confusion about what kube-proxy actually does (and doesn't do)?

kube-proxy handles Service-to-Pod routing within the cluster's data plane — it does not implement pod-to-pod networking itself; that's the CNI plugin's job (Calico, Cilium, Flannel). People sometimes debug pod connectivity issues by inspecting kube-proxy's iptables rules when the actual problem is CNI-level routing or network policy. Also, some CNIs like Cilium can replace kube-proxy entirely with eBPF-based routing, which is worth knowing before assuming kube-proxy is always in the traffic path.