Skip to main content

Kubernetes Cost Optimization Without Breaking SLOs

Average Kubernetes CPU utilization across production clusters is 8%. Here is the complete 2026 playbook for cutting cloud spend without touching your SLOs.

The Kubernetes bill arrives. It is double what was budgeted. Someone asks the engineering team what changed. Nothing specific changed — the cluster just grew because it was easier to over-provision than to tune. This is the default state of most Kubernetes clusters older than eighteen months.

The numbers are worse than most teams assume. Cast AI's 2026 State of Kubernetes Optimization report, drawn from tens of thousands of production clusters on AWS, Azure, and GCP, put average CPU utilization at 8%, memory at 20%, and GPU utilization at just 5% — with CPU overprovisioning up 69% year over year. For a company spending ₹50 lakh per month on cloud infrastructure, that gap is ₹35-45 lakh sitting idle every thirty days.

Why Kubernetes Clusters Overspend by Default

Kubernetes does not manage cost. It manages availability. When in doubt, Kubernetes does the safe thing: it keeps the resource alive, even if nothing is using it.

The four biggest cost drivers in a typical 2026 cluster:

  • Oversized resource requests: Developers set requests.memory: 2Gi and requests.cpu: 1000m out of caution. The actual usage is 200Mi and 80m. The scheduler reserves the full requested amount and nodes fill up with ghost capacity.
  • Idle workloads: Dev and staging environments run 24/7 even though nobody uses them between 10 PM and 9 AM.
  • Underutilized node groups: Fixed-size node groups provisioned for peak load sit at 15% utilization for 20 hours a day.
  • Unmanaged GPU/AI workload spend: As inference and training workloads move onto Kubernetes, pods increasingly claim a full GPU but use a fraction of it — GPU cost is now a fast-growing, frequently invisible line item.

Each of these has a specific fix, and none of them require touching your production SLOs.

Step 1: See What You Are Actually Using

You cannot optimize what you cannot measure. Start with resource usage visibility.

Bash
## Node-level CPU and memory usage
kubectl top nodes
## Pod-level usage
kubectl top pods --all-namespaces
## Find biggest resource consumers
kubectl top pods -A --sort-by=memory | head -20

For a proper view, deploy Goldilocks — it watches actual usage and recommends correct requests and limits:

Bash
helm repo add fairwinds-stable https://charts.fairwinds.com/stable
helm install goldilocks fairwinds-stable/goldilocks \
--namespace goldilocks \
--create-namespace
## Label a namespace to enable recommendations
kubectl label namespace production \
goldilocks.fairwinds.com/enabled=true

Open the Goldilocks dashboard and you will see a table showing every deployment's actual CPU/memory usage vs its requested values — and an auto-generated recommendation for what the requests should actually be. This is the same entry point Cast AI and other 2026 benchmarks recommend for teams just starting out, before committing to a paid autonomous platform.

Step 2: Fix Resource Requests

This single step is typically worth 30-40% cost reduction for teams that have never done it — and closer to 50%+ if your cluster matches the 2026 industry-average 8% CPU utilization figure above.

A pod with oversized requests blocks scheduler capacity even while idle. The node reports "full" to the scheduler when it is actually at 20% real utilization.

Here is the pattern: use VPA (Vertical Pod Autoscaler) in recommendation mode to generate correct values, then apply them to your manifests.

YAML
apiVersion: autoscaling.k8s.io/v1
kind: VerticalPodAutoscaler
metadata:
name: payment-service-vpa
namespace: production
spec:
targetRef:
apiVersion: apps/v1
kind: Deployment
name: payment-service
updatePolicy:
updateMode: "Off" ## recommendation only, no auto-apply
Bash
## View VPA recommendations after 24 hours of data
kubectl describe vpa payment-service-vpa -n production

Apply the recommended values to your Deployment. Then do this for every service. It is tedious — it is also the highest-ROI work you will do all quarter.

2026 update — in-place resize: Kubernetes v1.35 stabilized in-place pod resize, meaning requests/limits can now be updated without restarting the pod in most cases. This removes the biggest historical friction point in acting on VPA recommendations continuously instead of as a quarterly cleanup project — check your cluster version before relying on it in production.

Step 3: Right-Size Your Node Groups with Cluster Autoscaler

Static node groups are a cost trap. A node group provisioned for 100 pods during a Swiggy dinner-rush peak sits at 12 pods at 3 AM.

Cluster Autoscaler (CA) scales your node groups based on pending pods and removes underutilized nodes automatically.

YAML
## Cluster Autoscaler key configuration
autoDiscovery:
clusterName: prod-cluster
extraArgs:
scale-down-utilization-threshold: "0.5" ## remove nodes below 50% use
scale-down-delay-after-add: "10m" ## wait before scaling down
skip-nodes-with-local-storage: "false"
balance-similar-node-groups: "true"

Pair Cluster Autoscaler with Karpenter if you are on AWS — Karpenter provisions exact-fit instances rather than pre-defined node types, which eliminates the "I need 4 CPUs but the only node type available is 8" waste.

Bash
## Install Karpenter (AWS)
helm upgrade --install karpenter oci://public.ecr.aws/karpenter/karpenter \
--namespace karpenter \
--create-namespace \
--version "${KARPENTER_VERSION}"

Step 4: Scale Non-Production to Zero

Your staging cluster does not need to run at 2 AM on a Sunday. Dev environments do not need to run at all during weekends.

KEDA (Kubernetes Event-Driven Autoscaling) can scale deployments to zero on a schedule and scale them back up before business hours:

YAML
apiVersion: keda.sh/v1alpha1
kind: ScaledObject
metadata:
name: staging-api-scaledown
namespace: staging
spec:
scaleTargetRef:
name: api-service
minReplicaCount: 0 ## allows scale-to-zero
maxReplicaCount: 5
triggers:
- type: cron
metadata:
timezone: Asia/Kolkata
start: "30 8 * * 1-5" ## scale up 8:30 AM weekdays
end: "0 21 * * 1-5" ## scale down 9 PM weekdays
desiredReplicas: "3"

For dev namespaces, go further: use kube-downscaler to automatically scale everything to zero outside working hours across entire namespaces, not just individual deployments.

Step 5: Use Spot Instances for Stateless Workloads

Spot (AWS) / Preemptible (GCP) / Spot (Azure) instances cost 60-90% less than on-demand. For stateless workloads — API servers, workers, batch jobs — they are a direct cost lever.

The key is a correct pod disruption budget and fast restart behavior:

YAML
apiVersion: policy/v1
kind: PodDisruptionBudget
metadata:
name: api-pdb
namespace: production
spec:
minAvailable: 2 ## always keep 2 pods alive during node drain
selector:
matchLabels:
app: api-service

Pair this with a node taint for spot nodes so only workloads that tolerate interruption get scheduled there:

YAML
tolerations:
- key: "node.kubernetes.io/spot"
operator: "Exists"
effect: "NoSchedule"

Never run stateful workloads (databases, Redis with persistence, Kafka) on spot nodes.

Step 6: Don't Ignore GPU Waste

If your cluster runs any inference or training workloads, GPU is now frequently the single most expensive and least-utilized resource class — 2026 benchmark data puts average GPU utilization at just 5% across production clusters. Fractional GPU allocation (letting multiple pods share a GPU based on actual memory/compute need instead of claiming it whole) is quickly becoming a baseline expectation rather than a nice-to-have. If you have no visibility into per-pod GPU utilization today, that is the next gap to close after Steps 1-5.

The Dashboard You Need

Wire everything together in Grafana with a cost dashboard that shows spend by namespace, by team, and by workload. OpenCost is the open-source standard for this:

Bash
helm install opencost \
opencost/opencost \
--namespace opencost \
--create-namespace

OpenCost breaks down your cloud bill by Kubernetes namespace, label, and deployment — so when engineering leadership asks "which team is spending the most," you have a real answer in thirty seconds. If you need automated, continuous rightsizing rather than a dashboard, that is a step up to an autonomous optimization platform (Cast AI, ScaleOps, and similar) — but get the free, open-source visibility layer working first.

Production Implementation Guidelines

Don't optimize and change SLOs at the same time. Run two weeks of VPA recommendations in observation mode before applying anything to production. Compare actual resource usage at peak load (dinner time for Swiggy, market open for Zerodha, festive sale launch for Flipkart) to the VPA recommendation before trusting it.

Set [LimitRange](/glossary/kubernetes-limitrange) objects in every namespace to prevent new deployments from landing without resource requests:

YAML
apiVersion: v1
kind: LimitRange
metadata:
name: default-limits
namespace: production
spec:
limits:
- default:
memory: 512Mi
cpu: 500m
defaultRequest:
memory: 128Mi
cpu: 100m
type: Container

This ensures that even if a developer forgets to set requests, the namespace defaults kick in — preventing ghost capacity from accumulating silently.

Trade-offs and Alternatives

Technique Savings Potential Risk Level
Fix resource requests 30-50% Low
Scale non-prod to zero 15-25% Very low
Cluster Autoscaler 10-20% Low
Spot instances 40-70% Medium
Karpenter 20-35% Low
Fractional GPU allocation 30-60% (GPU spend only) Medium

Start with fixing resource requests and scaling non-prod to zero. These two alone typically justify the time investment in the first week.

Note

References and Further Reading

Frequently Asked Questions

What is a realistic Kubernetes CPU utilization target?

2026 industry benchmarks show average production CPU utilization around 8%, so even getting to 40-50% sustained utilization represents a major improvement over the norm — pushing much higher risks removing the headroom you need for traffic spikes.

Does in-place pod resize (Kubernetes v1.35) remove the need for VPA?

No — VPA still generates the recommendation. In-place resize just means you can apply that recommendation without a pod restart, which makes continuous rightsizing practical instead of a disruptive quarterly event.

Should I use an autonomous cost-optimization platform or stay with open-source tools?

Open-source (Goldilocks, VPA, Cluster Autoscaler, OpenCost) gets most teams 30-50% of the available savings for no license cost. Autonomous platforms add continuous, automatic enforcement on top — worth it once manual tuning stops keeping pace with cluster growth.

What's the fastest single change to reduce a bloated Kubernetes bill?

Fixing oversized resource requests — typically worth 30-40% cost reduction for teams that have never tuned them, since a pod requesting far more CPU/memory than it uses blocks scheduler capacity on a node even while the node is mostly idle.

Why is GPU utilization often the most overlooked cost driver on Kubernetes?

Because pods typically claim a full GPU even when using only a fraction of its capacity, and 2026 benchmark data puts average GPU utilization at just 5% across production clusters — without fractional GPU allocation or per-pod GPU visibility, that waste stays invisible on a standard cost dashboard.

Discussion0