Main content
DevOps

Idle GPUs Are the New Cloud Waste: Why Kubernetes GPU Utilization Averages 5%

Cast AI analyzed about 23,000 Kubernetes clusters and found average GPU utilization of 5%. Here is what that means for your bill and the practical steps that bring it down.

Illustration for “Idle GPUs Are the New Cloud Waste: Why Kubernetes GPU Utilization Averages 5%”

An idle CPU costs cents an hour. An idle GPU costs dollars an hour. That gap is why one statistic from Cast AI's 2026 report deserves attention from every team running AI on Kubernetes: across roughly 23,000 clusters analyzed, average GPU utilization was 5%, against 8% for CPU and 20% for memory. Put plainly, organizations were provisioning about twenty times more GPU capacity than they actually used.

The Math That Makes This Expensive

The effective price of a GPU is its hourly rate divided by how busy it is. The numbers below are an illustration, not a quote from any provider.

Utilization Hourly price Cost per hour of useful work
5%$10$200
20%$10$50
40%$10$25

Raising utilization is therefore worth more than almost any discount you can negotiate. Going from 5% to 40% is an eightfold improvement in cost per useful hour with no change to your hourly rate.

Why This Is Getting Worse, Not Better

Two broader signals point the same way. Flexera's 2026 State of the Cloud report estimates that 29% of IaaS and PaaS spend is wasted, the first increase in five years, with AI workloads cited as a driver. And capacity is getting pricier: in January 2026 AWS raised EC2 Capacity Blocks for ML prices for H200 instances by about 15%, as reported at the time. Idle capacity costs more each quarter, which makes the waste more visible on the bill. We cover the wider picture in cloud waste is back at 29%.

Where GPU Idle Time Comes From

  • Peak-sized reservations. Teams reserve GPUs for the busiest hour and pay for them all day.
  • Whole-GPU allocation. A small inference service or a notebook requests a full device and uses a fraction of it.
  • Always-on development environments. Notebooks and experiment nodes left running overnight and over weekends.
  • Stalled or queued training. Nodes stay up while jobs wait on data, checkpoints or a human.
  • Scheduling fragmentation. Pods that cannot be packed together leave partly used nodes that never scale down.
  • Forgotten endpoints. Hosted inference endpoints with zero traffic that nobody owns.

How to Raise GPU Utilization

1. Measure allocation against actual use

Most dashboards show GPUs requested. Add GPU utilization from your device metrics (for NVIDIA, the DCGM exporter) and track the ratio per namespace and per workload. Your first target is the list of workloads that request a GPU and barely use it.

2. Scale GPU node pools to zero

Configure autoscaling so empty GPU nodes are removed quickly, and let batch and development pools scale to zero when no pods are scheduled. This single change often removes the largest block of waste.

3. Share GPUs where workloads allow

NVIDIA Multi-Instance GPU (MIG) splits supported GPUs into isolated slices, and time-slicing lets several pods take turns on one device. Both suit inference services and notebooks that do not need a whole GPU. Training jobs that saturate a device are usually better left alone.

4. Put interruptible work on spot capacity

Training with regular checkpoints and batch inference tolerate interruptions and can run on discounted spot or preemptible capacity. Availability of the newest GPU types on spot is limited, so keep an on-demand fallback.

5. Schedule development GPUs

Shut down notebook and experiment GPUs outside working hours by default, and require an explicit extension rather than an explicit shutdown.

6. Commit last, not first

Buy reserved or committed GPU capacity only after you have fixed utilization. Committing to capacity you are barely using locks in the waste for one to three years. Our guide to Savings Plans, Reserved Instances and Spot explains how to size commitments.

A Quick GPU Audit You Can Run This Week

  1. List what is allocated. Running kubectl describe nodes on your GPU nodes shows allocatable and requested nvidia.com/gpu resources per node.
  2. Compare with what is used. If you run NVIDIA's DCGM exporter, graph DCGM_FI_DEV_GPU_UTIL per pod over two weeks. Workloads that request a GPU and rarely exceed single-digit utilization are your first candidates.
  3. Find GPU nodes with nothing scheduled. Any GPU node with zero GPU pods for a day or more should be draining and shutting down.
  4. Check managed endpoints. In SageMaker, Vertex AI and Azure ML, list endpoints and compute instances with no requests or sessions in the last seven days.
  5. Price the gap. Multiply idle hours by the hourly rate to get a monthly figure you can put in front of the owner.

Assign Ownership

GPU waste persists when no team sees the bill. Allocate GPU cost to the namespace or team that requests it, publish a weekly utilization report, and set a target such as a minimum average utilization for production inference. The allocation approach in our showback and chargeback guide works unchanged for GPUs. For Kubernetes-specific tuning beyond GPUs, see Kubernetes cost optimization.

How Varcio Helps

Varcio flags idle GPU nodes in Kubernetes clusters, and its AI Infrastructure view surfaces idle GPU-backed resources across Amazon SageMaker, Google Vertex AI and Azure ML, such as endpoints with zero invocations or traffic, with an estimated monthly saving per finding. That gives you a ranked list to act on instead of a utilization chart to interpret. Explore the platform capabilities or talk to our FinOps team about a GPU spend review.

Frequently asked questions

What is the average GPU utilization in Kubernetes?

Cast AI analyzed about 23,000 Kubernetes clusters for its 2026 report and found average GPU utilization of 5%, alongside 8% for CPU and 20% for memory. In other words, organizations were allocating roughly 20 times more GPU capacity than they actively used.

How do you calculate the real cost of an idle GPU?

Divide the hourly price by utilization. A GPU billed at $10 per hour that is busy 5% of the time costs $200 for every hour of useful work. At 40% utilization the same GPU costs $25 per useful hour.

Can you share a GPU between workloads?

Yes. NVIDIA Multi-Instance GPU (MIG) partitions supported GPUs into isolated slices, and time-slicing lets several pods share a GPU in turn. Both help when individual workloads do not need a whole device, such as inference and notebooks.

Does Varcio detect idle GPU resources?

Yes. Varcio flags idle GPU nodes in Kubernetes clusters and surfaces idle GPU-backed endpoints and compute across Amazon SageMaker, Google Vertex AI and Azure ML, with estimated monthly savings for each finding.

Turn this into savings on your own estate

Connect a cloud account with read-only access and see costed, ranked findings from the first scan — or talk to our FinOps team about a program.