Main content
AI & GenAI

FinOps for AI: How to Control LLM Token Spend Before It Controls Your Budget

98% of FinOps teams now manage AI spend, up from 31% two years ago. Here is a practical playbook for tracking token costs, cutting LLM inference bills, and tying AI spend to business value.

Illustration for “FinOps for AI: How to Control LLM Token Spend Before It Controls Your Budget”

For most engineering teams, AI went from a pilot line item to a top-five cloud cost in about eighteen months. The State of FinOps 2026 report puts numbers on it: 98% of FinOps teams now manage AI spend, up from 31% two years earlier, and AI cost management is the skill set practitioners most want to build. The report draws on 1,192 practitioners representing more than $83 billion in annual cloud spend, so this is no longer an edge case.

This guide covers the practical side: how to see AI spend in one place, which levers cut it without hurting quality, and how to connect cost to value so the conversation with finance is about return rather than invoices.

Why AI Spend Is Harder to Forecast Than Cloud Spend

Traditional cloud costs track capacity you provision. LLM costs track behavior: how long prompts are, how many steps an agent takes, how often users retry. One product change, such as adding retrieved documents to every prompt, can multiply cost per request overnight with no infrastructure change at all. Flexera's 2026 State of the Cloud report lists cost unpredictability among the top challenges of scaling AI workloads, cited by 30% of respondents, which matches what we see in practice.

Spend also lands in several places at once: direct API invoices from model providers, managed services on your cloud bill (Amazon Bedrock, Azure OpenAI, Google Vertex AI), and the GPU instances behind self-hosted models. Until those are combined, nobody can answer the simple question of what AI costs per feature.

Step 1: Put All AI Spend in One View

Pull AI costs from three sources into the same model as the rest of your cloud bill:

  • Managed AI services on your cloud invoice: Bedrock model invocations and provisioned throughput, Azure OpenAI deployments, Vertex AI endpoints.
  • Direct provider usage: API invoices and usage exports from model vendors, normalized to the same dimensions (provider, model, team, environment).
  • Compute under your control: SageMaker, Vertex and Azure ML endpoints, notebooks, and GPU nodes in Kubernetes. See our companion piece on idle GPU waste.

The FOCUS specification exists to make this kind of normalization less painful, and it is the reason a shared schema across clouds and AI vendors is now realistic.

Step 2: Allocate to Owners

Tag or label every AI workload with a team, a product and an environment, and require API keys or projects per product rather than one shared key. A single shared key is the most common reason AI costs cannot be attributed. The same tagging discipline from our showback and chargeback guide applies directly.

Step 3: Define Unit Economics

The State of FinOps 2026 data shows 49% of organizations now track unit economics, up nine points year over year, and AI is the reason many are starting. Pick one unit per AI feature and report it monthly:

  • Cost per resolved support conversation
  • Cost per generated report or document
  • Cost per 1,000 requests, or per active user per month

A feature that costs $0.40 per resolved ticket is a story you can defend when the human-handled ticket costs several dollars. A feature that costs $0.40 per click on a free tier is a different story. Without a unit, you cannot tell the two apart.

Step 4: Pull the Right Cost Levers

Lever What it does Quality risk
Prompt cachingReuses repeated prompt prefixes (system prompts, documents) at a fraction of the normal input priceNone
Batch processingRuns non-urgent jobs asynchronously at a discountNone (latency only)
Output and context limitsCaps response length and removes unused retrieved contextLow, test it
Model routingSends easy requests to smaller, cheaper modelsMedium, needs evals
Idle endpoint cleanupShuts down hosted endpoints with no trafficNone

Prompt caching. If your requests repeat a long prefix, caching is usually the highest-return change. Anthropic's prompt caching documentation prices cache reads at 0.1x the base input price on most models (some newer models go lower), with cache writes at 1.25x for the five-minute cache and 2x for the one-hour cache. Check your model's table, and check that your cache read counters are non-zero, because prompts under the minimum cacheable length are silently processed without caching.

Batch processing. Evaluation runs, document classification, nightly enrichment and backfills rarely need an instant answer. The Message Batches API cuts cost by 50%, and most batches finish in under an hour. Other major providers offer comparable asynchronous options; confirm current pricing in their documentation.

Model routing. Not every request needs your most capable model. Classification, extraction and short summaries often work on a smaller model at a fraction of the cost. Build an evaluation set of real requests first, measure quality per route, and only then shift traffic.

Idle endpoints. Hosted endpoints bill while they wait. A SageMaker or Vertex endpoint left up after a proof of concept keeps charging, and a GPU-backed one charges a lot. Alert on endpoints with zero invocations over a week.

Step 5: Add Guardrails

  • Budgets per product and environment with alerts at 50%, 80% and 100%, so a runaway agent loop shows up the same day.
  • Anomaly detection on spend per model and per key. Our anomaly detection guide covers thresholds and alert routing.
  • Pre-deployment review for expensive choices: provisioned throughput commitments, large GPU instance types, and long-context features.
  • Rate and spend limits in the application, such as per-user quotas, so one misbehaving client cannot burn a month of budget.

Step 6: Tie Cost to Value

The State of FinOps 2026 findings describe a shift from cost control to technology value: mature teams quantify AI value and influence technology selection instead of only trimming bills. In practice that means every AI feature carries three numbers on one page: what it costs per unit, what outcome it produces, and who owns it. Features that cannot name an outcome are candidates to cut, and features with strong unit economics earn a larger budget.

How Varcio Helps

The Varcio platform includes an AI Infrastructure view with resource-level findings for Amazon SageMaker, Amazon Bedrock, Google Vertex AI and Azure OpenAI / AI Foundry, and AI cost tracking that normalizes token, request and model spend across providers. It sits in the same workspace as the rest of your cloud cost, so AI shows up next to compute and storage rather than in a separate spreadsheet. See the full capability list, or talk to our AI and GenAI team about an architecture review.

A Four-Week Starter Plan

  • Week 1: Inventory every AI cost source and assign an owner to each key, project and endpoint.
  • Week 2: Turn on caching and batching where the workload allows; shut down idle endpoints.
  • Week 3: Define one unit metric per AI feature and publish the first report.
  • Week 4: Set budgets and anomaly alerts, and run a model-routing evaluation on your highest-volume feature.

Frequently asked questions

What is FinOps for AI?

FinOps for AI applies FinOps practices (visibility, allocation, optimization and accountability) to AI spend: LLM API tokens, managed model services such as Amazon Bedrock, Azure OpenAI and Google Vertex AI, and the GPU infrastructure behind training and inference. The goal is to make AI cost visible per team, product and feature, and to tie it to the value it produces.

How do you reduce LLM token costs without hurting quality?

Start with the levers that do not change answers: prompt caching for repeated context, batch processing for non-urgent work, output length limits, and trimming unused context. Then move to model routing, sending simple requests to smaller models and reserving frontier models for hard ones, and validate each change against an evaluation set before rolling it out.

What is a good unit metric for AI cost?

Choose a unit the business already understands: cost per resolved support ticket, per generated document, per 1,000 requests, or per active user. Raw token counts are useful for engineers but do not show whether a feature is economically viable.

Does Varcio track AI spend?

Yes. Varcio includes an AI Infrastructure view with resource-level findings for Amazon SageMaker, Amazon Bedrock, Google Vertex AI and Azure OpenAI / AI Foundry, and AI cost tracking that normalizes token, request and model spend across providers using FOCUS-compatible fields.

Turn this into savings on your own estate

Connect a cloud account with read-only access and see costed, ranked findings from the first scan — or talk to our FinOps team about a program.