FinOps Evolves to Manage Generative AI Spend

FinOps practitioners are adapting to generative AI's token-based pricing: the FinOps Foundation's State of FinOps 2026 survey found that 98% of teams now manage AI spend, up from just 31% two years ago, with granular AI cost monitoring the industry's most-requested FinOps tool feature. At FinOps X 2026, Fidelity Investments' Jennifer Hays told theCUBE that token costs ripple into a dozen or more adjacent expenses, from database throughput to developer laptops, while HSBC's Natalie Daley said the discipline must help choose the right model for the job, not just track spend. Separately, theCUBE Research found 24% of organizations now want to release code on an hourly basis.
FinOps teams are not just anecdotally worried about generative AI costs, the shift is now measured at industry scale: the FinOps Foundation's own 2026 survey of nearly 1,200 practitioners found the share managing AI spend jumped from 31% two years ago to 98% today, and "granular monitoring of AI spend (tokens, LLM requests, and GPU utilization)" is now the single most-requested FinOps tool capability.
What happened
At FinOps X 2026, Jennifer Hays, senior vice president and head of engineering excellence and technology strategy execution at Fidelity Investments, told theCUBE (SiliconANGLE Media's livestreaming studio, a disclosed paid media partner for the event) that token-based pricing creates a wide tail of adjacent costs: "You have to get transparency in your token costs, but you have to understand actually how it impacts probably a dozen or more costs around you," including database throughput, data volume into platforms like Snowflake, and even whether developers run models locally on their own hardware. Natalie Daley, HSBC's director and global head of cloud economics and FinOps, said the discipline's job is expanding from tracking cost to helping teams choose the right model for the right job. TheCUBE Research separately found that 24% of organizations want to release code on an hourly basis, a cadence Nashawaty linked to the pace of new model releases.
Financial context
The FinOps Foundation's State of FinOps 2026 report, based on 1,192 respondents representing more than $83 billion in annual cloud spend, confirms the scale of this shift independently of the conference commentary: AI is now the FinOps discipline's top forward-looking priority, and 98% of practitioners manage AI spend, up from 63% in 2025 and 31% in 2024. "Granular monitoring of AI spend (tokens, LLM requests, and GPU utilization)" ranks as the single most-requested tool feature, ahead of pre-deployment cost estimation and a unified cost dashboard.
For practitioners
The practical lesson echoed across both the vendor-neutral survey and the practitioner interviews is that token counts alone understate AI's true cost footprint. Teams should budget for adjacent costs, database and data-movement charges, developer hardware, and retrieval or caching infrastructure, alongside model inference, and should treat model selection itself, which model for which task, as a cost lever and not only a technical one.
What to watch
- •Whether FOCUS, the FinOps Open Cost and Usage Specification, gains broader AI-workload support, a top practitioner request in the 2026 survey.
- •Vendor tooling that adds token- and GPU-level granularity to cost dashboards.
- •Whether organizations reporting hourly release cadences (24% per theCUBE Research) see AI cost governance keep pace with deployment speed.
- •Further FinOps Foundation data on how AI cost management responsibilities are being staffed and budgeted.
Key Points
- 1FinOps Foundation survey data shows 98% of teams now manage AI spend, up from 31% two years ago, prioritizing granular token and GPU cost monitoring.
- 2Fidelity's Jennifer Hays told theCUBE that token costs ripple into a dozen or more adjacent expenses, including database throughput, data movement, and developer hardware.
- 3Practitioners should budget for costs beyond model inference and treat model selection itself as a cost lever, per HSBC's Natalie Daley.
Scoring Rationale
Base content originates from paid theCUBE/SiliconANGLE conference coverage of FinOps X, a category this audit typically discounts versus independent reporting; corroborating the practitioner anecdotes with the FinOps Foundation's own 2026 survey data (98% now manage AI spend, up from 31% in 2024) adds real substance, but the story remains routine practitioner commentary rather than a discrete news event.
Sources
Public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

