UBS Finds Enterprises Throttling AI Spending
UBS analysts led by Karl Keirstead found that roughly 60% of enterprises are now throttling AI spending in some way, based on conversations with about a dozen enterprise IT executives, according to a UBS research note relayed via Business Insider and Futu/Wall Street CN reporting in June 2026. For practitioners, this marks token-cost optimization as a mainstream IT priority rather than a niche concern: one surveyed company reported a single user running up $35,000 in monthly AWS Bedrock costs, and some teams exceeded weekly token quotas by up to 200%. UBS describes the trend as a modest "emerging headwind" for AI model makers rather than a stalled-adoption signal, and says enterprises increasingly use DeepSeek and other open or Chinese models via "model routing" to cut costs on simpler tasks while reserving premium models like Claude for complex work. The analysts said they "are not ringing the alarm bells" and called the shift "a healthy problem."
For teams building or buying AI infrastructure, this UBS research note is one of the clearest signals yet that token-cost governance has moved from a nice-to-have into a standard IT function - with direct implications for which models get default routing, how usage is metered, and which vendors face near-term revenue pressure.
What happened
According to a UBS research note dated June 23, 2026, covered by Business Insider and by Futu/Wall Street CN, UBS analysts Karl Keirstead, Timothy Arcuri, and Taylor McGinnis wrote that, based on conversations with roughly a dozen enterprise IT executives, about 60% of enterprises were "in some manner throttling AI spend" by adding guardrails such as token pooling, model downgrading, waste alerts, and per-user usage limits. Reported examples include one company that exhausted a large share of its annual token budget and cut its internal AI tools from five to two, one user who ran up $35,000 in a single month on AWS Bedrock, and DevOps teams that consistently exceeded weekly token quotas by 100-200%. The analysts called the pattern a modest "emerging headwind" for AI model makers, said "token spend optimization has become a key issue in most organizations," and added they "are not ringing the alarm bells," calling it "a healthy problem."
Technical context
The primary technical response enterprises are adopting is not blanket rate-limiting but model routing: directing routine tasks to cheaper models and reserving expensive frontier models for complex reasoning, code generation, or long-context work. Per Futu/Wall Street CN's reporting on the same UBS note, the price gap motivating this is real - citing Anthropic's public pricing as an example, Haiku costs a fraction of Opus- or Fable/Mythos-tier pricing per output token, making task-based model selection highly cost-effective. Engineering teams are pairing routing with familiar cost-control patterns: token pooling, batching, caching, context-window pruning, and client-side usage accounting.
Industry context
The note names open-source and Chinese models - including DeepSeek, Alibaba's Qwen, MiniMax, and Zhipu AI's GLM - as beneficiaries of this shift, with one large global bank reportedly deploying Qwen on-premises to offset spending on premium models like Claude. UBS frames this as normal cost governance during early-stage enterprise AI adoption rather than a sign that AI adoption itself is slowing; a companion UBS survey of roughly 130 companies found only 8% have deployed AI agents at scale in production, suggesting the token-consumption growth this optimization responds to is still in its early stages.
For practitioners
Engineers and ML platform teams should treat this as validation for investing now in routing infrastructure, per-task cost accounting, and usage observability rather than waiting for bills to force the issue - vendors and internal tooling that make routing and cost attribution easy are positioned to benefit regardless of which model wins any given task.
What to watch
- •Whether AI model vendors respond with tiered pricing or bundled inference options to blunt the shift toward cheaper models.
- •Adoption of token-level observability and chargeback tooling inside enterprises.
- •Procurement shifts toward smaller or fine-tuned models for high-volume, low-complexity tasks.
Key Points
- 1UBS found about 60% of enterprises are now throttling AI spend, based on conversations with roughly a dozen enterprise IT executives in June 2026.
- 2Enterprises are using model routing to shift routine tasks to cheaper or open-source models like DeepSeek and Qwen while reserving premium models for complex work.
- 3UBS calls this a modest headwind, not stalled adoption; a companion survey found only 8% of companies deploy AI agents at scale in production.
Scoring Rationale
A well-corroborated (verified independently against the underlying UBS note via two separate reprints plus CNBC/AI-Insider) signal of a real shift in enterprise AI purchasing behavior, with concrete figures (60% throttling, $35K/month outlier, 8% agent-at-scale) rather than vague trend talk. Directly relevant to model vendors and ML platform teams; held below 'major' since UBS itself frames it as a modest, expected headwind rather than a structural break.
Sources
Primary source and supporting public references used for this report.
View 3 more sources
Practice with real SaaS & B2B data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all SaaS & B2B problems
