On this page · 8 sections
An idle CPU costs cents an hour. An idle GPU costs dollars an hour. That gap is why one statistic from Cast AI's 2026 report deserves attention from every team running AI on Kubernetes: across roughly 23,000 clusters analyzed, average GPU utilization was 5%, against 8% for CPU and 20% for memory. Put plainly, organizations were provisioning about twenty times more GPU capacity than they actually used.
The Math That Makes This Expensive
The effective price of a GPU is its hourly rate divided by how busy it is. The numbers below are an illustration, not a quote from any provider.
| Utilization | Hourly price | Cost per hour of useful work |
|---|---|---|
| 5% | $10 | $200 |
| 20% | $10 | $50 |
| 40% | $10 | $25 |
Raising utilization is therefore worth more than almost any discount you can negotiate. Going from 5% to 40% is an eightfold improvement in cost per useful hour with no change to your hourly rate.
Why This Is Getting Worse, Not Better
Two broader signals point the same way. Flexera's 2026 State of the Cloud report estimates that 29% of IaaS and PaaS spend is wasted, the first increase in five years, with AI workloads cited as a driver. And capacity is getting pricier: in January 2026 AWS raised EC2 Capacity Blocks for ML prices for H200 instances by about 15%, as reported at the time. Idle capacity costs more each quarter, which makes the waste more visible on the bill. We cover the wider picture in cloud waste is back at 29%.
Where GPU Idle Time Comes From
- Peak-sized reservations. Teams reserve GPUs for the busiest hour and pay for them all day.
- Whole-GPU allocation. A small inference service or a notebook requests a full device and uses a fraction of it.
- Always-on development environments. Notebooks and experiment nodes left running overnight and over weekends.
- Stalled or queued training. Nodes stay up while jobs wait on data, checkpoints or a human.
- Scheduling fragmentation. Pods that cannot be packed together leave partly used nodes that never scale down.
- Forgotten endpoints. Hosted inference endpoints with zero traffic that nobody owns.
How to Raise GPU Utilization
1. Measure allocation against actual use
Most dashboards show GPUs requested. Add GPU utilization from your device metrics (for NVIDIA, the DCGM exporter) and track the ratio per namespace and per workload. Your first target is the list of workloads that request a GPU and barely use it.
2. Scale GPU node pools to zero
Configure autoscaling so empty GPU nodes are removed quickly, and let batch and development pools scale to zero when no pods are scheduled. This single change often removes the largest block of waste.
3. Share GPUs where workloads allow
NVIDIA Multi-Instance GPU (MIG) splits supported GPUs into isolated slices, and time-slicing lets several pods take turns on one device. Both suit inference services and notebooks that do not need a whole GPU. Training jobs that saturate a device are usually better left alone.
4. Put interruptible work on spot capacity
Training with regular checkpoints and batch inference tolerate interruptions and can run on discounted spot or preemptible capacity. Availability of the newest GPU types on spot is limited, so keep an on-demand fallback.
5. Schedule development GPUs
Shut down notebook and experiment GPUs outside working hours by default, and require an explicit extension rather than an explicit shutdown.
6. Commit last, not first
Buy reserved or committed GPU capacity only after you have fixed utilization. Committing to capacity you are barely using locks in the waste for one to three years. Our guide to Savings Plans, Reserved Instances and Spot explains how to size commitments.
A Quick GPU Audit You Can Run This Week
- List what is allocated. Running
kubectl describe nodeson your GPU nodes shows allocatable and requestednvidia.com/gpuresources per node. - Compare with what is used. If you run NVIDIA's DCGM exporter, graph
DCGM_FI_DEV_GPU_UTILper pod over two weeks. Workloads that request a GPU and rarely exceed single-digit utilization are your first candidates. - Find GPU nodes with nothing scheduled. Any GPU node with zero GPU pods for a day or more should be draining and shutting down.
- Check managed endpoints. In SageMaker, Vertex AI and Azure ML, list endpoints and compute instances with no requests or sessions in the last seven days.
- Price the gap. Multiply idle hours by the hourly rate to get a monthly figure you can put in front of the owner.
Assign Ownership
GPU waste persists when no team sees the bill. Allocate GPU cost to the namespace or team that requests it, publish a weekly utilization report, and set a target such as a minimum average utilization for production inference. The allocation approach in our showback and chargeback guide works unchanged for GPUs. For Kubernetes-specific tuning beyond GPUs, see Kubernetes cost optimization.
How Varcio Helps
Varcio flags idle GPU nodes in Kubernetes clusters, and its AI Infrastructure view surfaces idle GPU-backed resources across Amazon SageMaker, Google Vertex AI and Azure ML, such as endpoints with zero invocations or traffic, with an estimated monthly saving per finding. That gives you a ranked list to act on instead of a utilization chart to interpret. Explore the platform capabilities or talk to our FinOps team about a GPU spend review.
Frequently asked questions
What is the average GPU utilization in Kubernetes?
Cast AI analyzed about 23,000 Kubernetes clusters for its 2026 report and found average GPU utilization of 5%, alongside 8% for CPU and 20% for memory. In other words, organizations were allocating roughly 20 times more GPU capacity than they actively used.
How do you calculate the real cost of an idle GPU?
Divide the hourly price by utilization. A GPU billed at $10 per hour that is busy 5% of the time costs $200 for every hour of useful work. At 40% utilization the same GPU costs $25 per useful hour.
Can you share a GPU between workloads?
Yes. NVIDIA Multi-Instance GPU (MIG) partitions supported GPUs into isolated slices, and time-slicing lets several pods share a GPU in turn. Both help when individual workloads do not need a whole device, such as inference and notebooks.
Does Varcio detect idle GPU resources?
Yes. Varcio flags idle GPU nodes in Kubernetes clusters and surfaces idle GPU-backed endpoints and compute across Amazon SageMaker, Google Vertex AI and Azure ML, with estimated monthly savings for each finding.