AI Models Neutral 6

Hundreds of billions in LLMs, yet AI pricing is a black box for enterprises

The AI industry’s token-based economics make pricing LLM-powered services a guessing game. Subtle prompt variations and agentic frameworks cause unpredictable token consumption, challenging enterprise adoption.

· 4 min read · Verified by 2 sources ·
Share

Key Takeaways

  • The AI industry’s token-based economics make pricing LLM-powered services a guessing game.
  • Subtle prompt variations and agentic frameworks cause unpredictable token consumption, challenging enterprise adoption.

Mentioned

Microsoft company MSFT Google company GOOGL Anthropic company Saviynt company Simon Gooch person Goldman Sachs company GS ChatGPT product Claude product Gemini product Artificial Intelligence technology

Key Intelligence

Key Facts

  1. 1Microsoft, Google, and Anthropic have collectively invested hundreds of billions of dollars in developing LLMs, with total industry spending still accelerating.
  2. 2According to Goldman Sachs, per-token costs have sharply declined, yet overall token consumption by businesses has surged, making total costs unpredictable.
  3. 3Agentic AI systems, which use multiple AI agents in coordination, dramatically increase token usage compared to single-prompt interactions.
  4. 4Saviynt’s Simon Gooch warns that fixed-cost, multi-year pricing models for AI services are impractical because token economics change too rapidly.
  5. 5Subtle variations in a user’s prompt can produce significantly different token-consuming responses, undermining accurate cost estimation.
  6. 6Free tiers of ChatGPT, Claude, and Gemini have set a consumer price anchor of zero, putting pressure on paid AI pricing strategies.

Trying to tie someone into a cost model for the next 12 months, two years, three years, it doesn't make any sense, honestly, because we don't know.

Simon Gooch at Saviynt

On AI pricing challenges

Analysis

For AI researchers and ML engineers, the pricing dilemma reveals a deeper flaw: current LLM architectures produce inherently stochastic outputs, making cost prediction almost impossible without rigorous guardrails. This unpredictability threatens to slow the rollout of mission-critical AI agents in enterprise settings.

The scramble to monetize artificial intelligence is hitting a wall: the economics of token-based pricing are so chaotic that even the industry's biggest players can't create stable cost models. As Microsoft, Google, and Anthropic collectively pour hundreds of billions into large language models (LLMs), the task of converting that spending into predictable revenue streams is proving fraught.

As a result, the same business workflow could cost $10 one day and $100 the next, depending on how the AI interprets the task and invokes sub-agents.

At the heart of the problem lies the token — the mathematical subunit that LLMs use to process prompts and generate responses. When a user asks ChatGPT to draft an email or an agentic system orchestrates a multi-step workflow, every input and output is broken into tokens. The bill for that interaction depends on how many tokens are consumed, which varies dramatically. A seemingly minor rephrase of a question can double or triple the token count; the same prompt can yield different-length answers from different models; and as noted by Saviynt’s Simon Gooch, agentic setups where multiple AI agents collaborate compound this variability unpredictably.

Goldman Sachs’ analysis highlights a paradox: while per-token costs have plummeted over recent years, total token consumption by enterprises has surged, making aggregate costs harder to forecast. For AI service providers, this means the unit economics that underpin pricing are constantly shifting. A fixed-price contract for 12 or 24 months, once standard in SaaS and cloud services, becomes a gamble. “Trying to tie someone into a cost model for the next 12 months, two years, three years, it doesn't make any sense, honestly, because we don't know,” Gooch said.

The implications ripple across the AI ecosystem. For SaaS companies building on top of LLMs, the inability to price reliably threatens margins and slows go-to-market strategies. When a startup cannot confidently quote a client a monthly fee for an AI-powered analytics tool, the sales cycle stalls. For enterprise buyers, sticker shock from variable monthly bills under consumption-based models can kill adoption. CIOs used to predictable cloud invoices are wary of AI line items that swing by thousands of dollars overnight.

This pricing fog is more than a temporary friction — it reflects fundamental architectural traits of current LLMs. The stochastic nature of generative AI means outputs aren't deterministic, and agentic frameworks, which increasingly chain model calls for reasoning and action, amplify token usage exponentially. As a result, the same business workflow could cost $10 one day and $100 the next, depending on how the AI interprets the task and invokes sub-agents. The industry has responded with attempts to set usage caps and credits, but these are blunt instruments that often frustrate power users.

The three tech giants — Microsoft, Google, and Anthropic — each pursue different monetization paths, from Copilot subscriptions to per-token APIs to enterprise deals, yet none have cracked the predictability code. The pressure is likely to accelerate consolidation or push toward end-to-end platforms that control the entire stack, from model inference to application, to better forecast costs. Meanwhile, the existence of free, high-quality AI like ChatGPT sets a consumer price anchor of zero, making it even harder for third-party developers to charge a premium.

What to Watch

Looking ahead, several developments could reshape the tokenomic landscape. Model efficiency improvements — such as smaller, specialized models or token-optimization techniques — could reduce per-call token counts, but they won't eliminate variability. More likely, pricing will bifurcate: mission-critical agentic workflows will adopt outcome-based models, where businesses pay for results achieved rather than raw token consumption, while simpler generative features remain on per-token or subscription plans. The rise of open-source LLMs could also apply downward price pressure, but exacerbate unpredictability as organizations mix models with different token behaviors.

The challenge identified by Saviynt’s Gooch underscores that the AI industry’s pricing immaturity is a strategic threat. Without transparent, predictable pricing, the hundreds of billions in investment risk being stranded as enterprises hesitate to commit. The winners will be those that build robust monitoring tools to forecast token usage and offer hybrid pricing that blends predictability with value. Until then, AI pricing remains a high-stakes experiment.

Sources

Sources

Based on 2 source articles

Cite This Page

"Hundreds of billions in LLMs, yet AI pricing is a black box for enterprises." AI Intelligence Brief, August 4, 2026. https://getaibrief.com/story/ai-pricing-token-black-box

How we covered this story

Every story in our AI coverage is assembled from multiple primary sources, cross-referenced for factual consistency, and scored along three independent dimensions: sentiment, operational impact, and source-cluster confidence. Single-source rumors and unverifiable claims do not pass our editorial gate. When a story shows "Verified by N sources" with N≥2, the development is independently corroborated; when N=1, we mark it explicitly so readers can weigh the signal accordingly.

Impact scoring uses a 1-10 scale weighted toward regulatory, financial, and operational consequence rather than coverage volume. A topic that runs in every outlet but moves no real decisions ranks lower than a niche regulatory filing that reshapes how operators in the AI space have to behave. Read our full methodology for the scoring rubric, our glossary for term definitions, and our trends index for the longitudinal view across the beat.

Sources are only linked to a story once they clear our classification pipeline at a minimum 35 percent relevance threshold. According to that methodology, reviewed July 2026, this follows multi-source corroboration standards recommended by journalism research bodies such as the Reuters Institute for the Study of Journalism.

See something wrong in this story — a wrong fact, a broken source link, a misattributed entity? Report a data issue.