Kubernetes made deployment easier and cost management harder. Your cloud bill shows EC2 instances or VMs, not the 40 teams sharing them. Engineers set generous resource requests "just in case." Clusters scale up and rarely scale back down.
This guide explains Kubernetes cost optimization step by step for EKS, AKS and GKE.
Kubernetes made deployment easier and cost management harder.
Why are Kubernetes costs hard to manage?
Three reasons:
- Shared infrastructure. Many teams share the same nodes, so cloud tags cannot show who used what.
- Requests drive cost, not usage. The scheduler reserves capacity based on requests. A pod requesting 2 CPUs but using 0.2 still blocks 2 CPUs.
- Layers of abstraction. Nodes, node groups, pods, namespaces and persistent volumes all contribute to cost in different ways.
How do you allocate Kubernetes costs to teams?
Use a Kubernetes cost allocation tool such as OpenCost (a CNCF project) or Kubecost. These tools combine cloud pricing with cluster metrics to calculate cost per namespace, deployment, label or pod.
Kubernetes cost allocation by namespace and label
Illustration in progress
Make namespaces and labels match your cloud tagging schema, for example team, application and environment. Then decide how to allocate shared costs such as idle capacity, system namespaces and control plane fees. See our cost allocation guide.
How do you reduce Kubernetes costs?
Follow this order, from highest impact to fine-tuning:
Kubernetes cost optimization levers
Illustration in progress
- Rightsize requests and limits. Compare requested CPU and memory with actual usage over two to four weeks. The Vertical Pod Autoscaler in recommendation mode can suggest values.
- Enable Horizontal Pod Autoscaler (HPA) so replicas follow demand rather than staying sized for peak.
- Use cluster autoscaling. Cluster Autoscaler or Karpenter (on EKS) add and remove nodes as pods need them. Karpenter can also choose cheaper instance types automatically.
- Improve bin-packing. Use a mix of node sizes and let the scheduler consolidate workloads onto fewer nodes.
- Run interruptible workloads on Spot. Stateless services, batch jobs and CI runners are good candidates. Use pod disruption budgets and multiple instance types.
- Scale down non-production clusters outside working hours.
- Clean up orphaned volumes and load balancers. Deleted workloads can leave persistent volumes and cloud load balancers running.
- Cover the baseline with commitments. After rightsizing, use Savings Plans or reservations for steady node capacity.
Requests vs usage: what good looks like
| Metric | Warning sign | Healthy target |
|---|---|---|
| CPU request utilisation | Under 30% | 50–70% |
| Memory request utilisation | Under 40% | 60–80% |
| Cluster idle cost | Over 30% of cluster spend | Under 15–20% |
| Spot share (eligible workloads) | 0% | Majority of stateless and batch |
These are general guidelines. Latency-sensitive or bursty workloads may need more headroom.
How does Kubernetes cost relate to observability?
You cannot rightsize what you cannot measure. Prometheus metrics for container CPU and memory are the input to every rightsizing decision. Teams that already run a good Kubernetes observability stack find cost optimisation much faster.
How Crozaint approaches Kubernetes costs
Crozaint runs Kubernetes on AWS, Azure and Google Cloud, as well as OpenShift, for clients across retail, financial services and technology. Our FinOps engagements bring cluster-level allocation into the same view as the rest of your cloud bill, so a team sees its namespaces and its managed services together.
Because we also deliver observability and 24/7 managed operations, we can rightsize with confidence: performance and reliability data sit next to cost data, so savings do not come at the expense of SLOs.
Common mistakes to avoid
- Copying the same large requests into every deployment
- Setting CPU limits that cause throttling, then raising requests to compensate
- Running stateful databases on Spot without a plan
- Ignoring idle cluster capacity in cost reports
- Optimising cost without watching latency and error rates
Conclusion
Kubernetes cost optimisation is mostly about honest requests, elastic scaling and clear ownership. Get visibility by namespace, fix requests, then automate scaling and Spot.
Running Kubernetes at scale? Book a 30-minute discovery call and we will review your cluster costs with you.

