Dive Brief:
- As enterprise AI use evolves from assistive features to multistep processes, tech leaders face growing inference bills, according to a Gartner report publish Monday. The analyst firm expects inference costs to increase more than fivefold through 2028.
- Tokens are becoming more cost efficient, but changing AI pricing models and more complex workflows are posing challenges for tech leaders in charge of cost management. The rate of innovation is currently outpacing the cost savings enterprises experience from their AI providers, per the report.
- Product leaders won’t be able to rely on more efficient tokens to keep costs in check for their organizations, Will Sommer, senior director analyst at Gartner, said in the report. “Each successive generation of AI capability will necessitate more, and often more expensive, tokens,” he said. “There is no reliable, economical one-size-fits-all model on the horizon.”
Dive Insight:
Cloud providers are increasing capital investments in AI infrastructure, despite enterprise cost concerns. Spending on compute to support model training and operation is expected to nearly double through the end of 2026, reaching $42 billion, Gartner found earlier this month.
The firm expects global token consumption to approach 300 trillion per day, an annualized growth rate of 500% to 600%, Sommer told CIO Dive in July.
Hard-to-predict AI costs are forcing some enterprises to rethink implementation plans. Nearly half of organizations reported that they’ve escalated AI spending surprises to the board, according to a July report by cost management software vendor Mavvrik.
Many organizations are tapping IT services providers for strategy support.
“Successfully implementing AI is hard," Sommer said in an email. "Effectively deploying and managing enterprise-wide agentic systems requires technologies and skillsets that most companies do not have today. For many enterprises, working in coordination with channel partners will be essential for aligning AI capabilities to business workflows, scaling deployments, maximizing ROI, and developing long-term AI strategies.”
Three factors are shaping token economics, according to Gartner. Foundational model costs are going down — a win for enterprises — but improved AI efficiency is unlocking more powerful, expensive models to chase higher-value AI applications. Those sophisticated AI workflows use far more tokens than earlier models’ simple chatbot interactions, which is driving up overall inference costs.
In a dynamic Gartner calls the "Inference Paradox," AI costs escalate without providing a clear pathway to predictable value.
“The harsh economics of the Inference Paradox are exemplified by the differences between a simple chatbot and an AI agent,” Sommer said in the report. “Where a simple chatbot must read and interpret a query and quickly respond with a probabilistically reasonable answer, an AI agent must constantly reason, negotiate and question itself.”
Achieving ROI from more advanced systems like agentic AI models or reasoning agents requires much higher returns than basic models, Gartner found. Otherwise, companies will need to highly optimize their models to complete complex tasks relative to cost-effective intelligence, Sommer said.
Both options are possible, but require significant overhauling of business workflows, Sommer said in July.
“Every CIO is going to have different priorities, they're going to have different things that are mission-critical to the organization,” Sommer said. “The CIO has to work with the rest of the organization to determine what the priorities are, and whether or not AI is a cost efficient way to solve those problems.”
Editor's note: Matt Ashare contributed reporting to this story.