Skip to main content

Toleration

A pod-level configuration in Kubernetes that allows the scheduler to place a pod onto a node that carries a matching Taint, enabling specific workloads to run on dedicated or restricted nodes.

What a toleration does

A toleration is the pod's way of saying "I am allowed on that restricted node." It is the matching key for a taint on a node. Without a matching toleration, the scheduler keeps the pod away from any node that carries the taint.

%%{init: {'themeVariables': {'fontSize': '18px'}, 'flowchart': {'nodeSpacing': 40, 'rankSpacing': 45, 'padding': 14}}}%% flowchart TD P1["Regular pod
no toleration"] -.-> N P2["GPU pod
tolerates workload=gpu"] ==> N N["Node gpu-node-1
taint workload=gpu:NoSchedule"] class P1,P2 base class N key classDef base fill:#f8fafc,stroke:#94a3b8,stroke-width:1.5px,color:#0f172a classDef key fill:#dbeafe,stroke:#2563eb,stroke-width:2px,color:#0f172a

Example

YAML
apiVersion: v1
kind: Pod
metadata:
name: video-encoder
spec:
tolerations:
- key: "workload"
operator: "Equal"
value: "gpu"
effect: "NoSchedule" # Must match the taint's effect
containers:
- name: encoder
image: registry.example.com/video-encoder:v3.1.0
resources:
limits:
nvidia.com/gpu: 1

The node side of this setup is a taint:

Bash
kubectl taint nodes gpu-node-1 workload=gpu:NoSchedule

Operators: Equal and Exists

  • Equal (the default) requires both the key and the value to match the taint.
  • Exists matches on the key alone and ignores the value.
  • An Exists toleration with no key matches every taint. This is a wildcard and should be limited to system DaemonSets such as log collectors.
YAML
tolerations:
- key: "workload"
operator: "Exists"
effect: "NoSchedule"

Taint effects and tolerationSeconds

Effect Behaviour for pods without a matching toleration
NoSchedule New pods are not scheduled on the node
PreferNoSchedule The scheduler avoids the node but may still use it
NoExecute New pods are blocked and running pods are evicted

For NoExecute taints, tolerationSeconds sets how long a pod may stay before it is evicted. Kubernetes uses this automatically for node failures: pods get a default 300-second toleration for the not-ready and unreachable taints, which you can override.

YAML
tolerations:
- key: "node.kubernetes.io/not-ready"
operator: "Exists"
effect: "NoExecute"
tolerationSeconds: 60 # Evict after 60 seconds instead of 300

Pairing with node affinity

A toleration only permits scheduling on a tainted node. It does not attract the pod there. To dedicate nodes properly, taint the nodes, add the toleration to the pods, and add a nodeSelector or nodeAffinity so the pods only land on those nodes.

YAML
spec:
tolerations:
- key: "workload"
operator: "Equal"
value: "gpu"
effect: "NoSchedule"
nodeSelector:
workload: gpu

Troubleshooting

Problem Symptom Fix
Pod stuck in Pending node(s) had untolerated taint Add a toleration with the matching key, value, and effect
Toleration set but pod still Pending Effect or value mismatch Make the toleration match the taint exactly
Pod lands on untainted nodes Toleration does not pin the pod Add nodeSelector or nodeAffinity
Non-GPU pods on expensive nodes Wildcard toleration in use Replace the empty-key Exists with a scoped key
Common Mistake

Using operator: Exists without a key. It tolerates every taint in the cluster and defeats taint-based isolation for that pod.

Security

For workloads that need strict isolation, such as PCI-scoped payment pods, use operator: Equal with the exact key and value, and combine it with node affinity so the pods run only on the dedicated nodes.

Frequently Asked Questions

How does a toleration actually interact with a taint to get a pod scheduled?

A taint on a node has three parts, key, value, and effect (NoSchedule, PreferNoSchedule, or NoExecute), and by default it repels every pod. A toleration on the pod spec must match the taint's key, value, and effect (or use the `Exists` operator to match any value) for the scheduler to consider that node eligible. Critically, a toleration only permits scheduling; it doesn't attract the pod there. Combine it with `nodeAffinity` or `nodeSelector` if you want the pod to go to that node.

What's a common misunderstanding about tolerations and NoExecute taints?

Engineers often assume a toleration is permanent, but a NoExecute toleration can specify `tolerationSeconds`, after which the pod is evicted even though it originally tolerated the taint. This matters for the not-ready and unreachable taints that Kubernetes adds automatically during outages: pods get a default of 300 seconds unless you set it explicitly, which affects failover timing.