Toleration
A pod-level configuration in Kubernetes that allows the scheduler to place a pod onto a node that carries a matching Taint, enabling specific workloads to run on dedicated or restricted nodes.
What a toleration does
A toleration is the pod's way of saying "I am allowed on that restricted node." It is the matching key for a taint on a node. Without a matching toleration, the scheduler keeps the pod away from any node that carries the taint.
no toleration"] -.-> N P2["GPU pod
tolerates workload=gpu"] ==> N N["Node gpu-node-1
taint workload=gpu:NoSchedule"] class P1,P2 base class N key classDef base fill:#f8fafc,stroke:#94a3b8,stroke-width:1.5px,color:#0f172a classDef key fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#0f172a
Example
apiVersion: v1kind: Podmetadata: name: video-encoderspec: tolerations: - key: "workload" operator: "Equal" value: "gpu" effect: "NoSchedule" # Must match the taint's effect containers: - name: encoder image: registry.example.com/video-encoder:v3.1.0 resources: limits: nvidia.com/gpu: 1The node side of this setup is a taint:
kubectl taint nodes gpu-node-1 workload=gpu:NoScheduleOperators: Equal and Exists
- Equal (the default) requires both the key and the value to match the taint.
- Exists matches on the key alone and ignores the value.
- An
Existstoleration with no key matches every taint. This is a wildcard and should be limited to system DaemonSets such as log collectors.
tolerations: - key: "workload" operator: "Exists" effect: "NoSchedule"Taint effects and tolerationSeconds
| Effect | Behaviour for pods without a matching toleration |
|---|---|
| NoSchedule | New pods are not scheduled on the node |
| PreferNoSchedule | The scheduler avoids the node but may still use it |
| NoExecute | New pods are blocked and running pods are evicted |
For NoExecute taints, tolerationSeconds sets how long a pod may stay before it is evicted. Kubernetes uses this automatically for node failures: pods get a default 300-second toleration for the not-ready and unreachable taints, which you can override.
tolerations: - key: "node.kubernetes.io/not-ready" operator: "Exists" effect: "NoExecute" tolerationSeconds: 60 # Evict after 60 seconds instead of 300Pairing with node affinity
A toleration only permits scheduling on a tainted node. It does not attract the pod there. To dedicate nodes properly, taint the nodes, add the toleration to the pods, and add a nodeSelector or nodeAffinity so the pods only land on those nodes.
spec: tolerations: - key: "workload" operator: "Equal" value: "gpu" effect: "NoSchedule" nodeSelector: workload: gpuTroubleshooting
| Problem | Symptom | Fix |
|---|---|---|
| Pod stuck in Pending | node(s) had untolerated taint |
Add a toleration with the matching key, value, and effect |
| Toleration set but pod still Pending | Effect or value mismatch | Make the toleration match the taint exactly |
| Pod lands on untainted nodes | Toleration does not pin the pod | Add nodeSelector or nodeAffinity |
| Non-GPU pods on expensive nodes | Wildcard toleration in use | Replace the empty-key Exists with a scoped key |
Common MistakeUsing
operator: Existswithout akey. It tolerates every taint in the cluster and defeats taint-based isolation for that pod.
SecurityFor workloads that need strict isolation, such as PCI-scoped payment pods, use
operator: Equalwith the exact key and value, and combine it with node affinity so the pods run only on the dedicated nodes.
Frequently Asked Questions
How does a toleration actually interact with a taint to get a pod scheduled?
A taint on a node has three parts, key, value, and effect (NoSchedule, PreferNoSchedule, or NoExecute), and by default it repels every pod. A toleration on the pod spec must match the taint's key, value, and effect (or use the `Exists` operator to match any value) for the scheduler to consider that node eligible. Critically, a toleration only permits scheduling; it doesn't attract the pod there. Combine it with `nodeAffinity` or `nodeSelector` if you want the pod to go to that node.
What's a common misunderstanding about tolerations and NoExecute taints?
Engineers often assume a toleration is permanent, but a NoExecute toleration can specify `tolerationSeconds`, after which the pod is evicted even though it originally tolerated the taint. This matters for the not-ready and unreachable taints that Kubernetes adds automatically during outages: pods get a default of 300 seconds unless you set it explicitly, which affects failover timing.