Agentic AI workloads can consume 10 to 100 times more tokens per task than conventional inference, putting pressure on enterprises to rethink per-token pricing as AI applications move into sustained production, according to a new Futurum Research report sponsored by QumulusAI (Nasdaq: QMLS).
The report, The Off Ramp From Per-Token Pricing: How Enterprises Regain AI Cost Control With Reserved Bare Metal, examines how the shift from simple prompt-and-response applications to multistep agentic systems is changing the economics of AI inference.
Futurum estimates that global inference spending will rise from $120 billion in 2025 to $885 billion by 2030, while agent and reasoning inference is expected to grow 219% in 2026. The report argues that the rising volume and complexity of inference workloads could accelerate enterprise adoption of reserved infrastructure alongside public cloud services.
Token Consumption Changes the Cost Equation
Traditional AI applications often make relatively straightforward inference calls. Agentic systems can perform multiple steps, invoke tools, retrieve information, evaluate outputs and make additional model calls before completing a task.
That changes the amount of compute required for a single business operation. Futurum says the per-task token footprint of agentic workloads can rise 10 to 100 times compared with a simple inference call. Under a pricing model tied directly to token consumption, that additional activity translates into a larger and potentially less predictable operating bill.
The report does not suggest that per-token services will disappear. Instead, it describes a progression in which companies can use serverless APIs and on-demand GPU infrastructure during experimentation and early deployment, then move workloads to reserved capacity once utilization becomes predictable.
Hybrid Infrastructure Gains Ground
Futurum’s survey of 824 AI decision-makers indicates that enterprises are already distributing AI workloads across multiple infrastructure environments.
Public-cloud infrastructure accounts for 41% of primary AI workload deployment, while organizations’ own data centers account for 36%. Colocation providers represent 13%, bare-metal and high-performance computing providers 6%, and edge infrastructure 5%. Taken together, 59% of respondents primarily run AI workloads outside hyperscaler public clouds.
The consumption picture is also weighted toward committed infrastructure. Reserved and owned accelerators account for 66% of AI compute consumption, compared with 19% for on-demand cloud, according to Futurum’s first-half 2026 survey.
The figures point to a market in which enterprises are combining infrastructure models rather than moving all workloads into one environment.
From Experimentation to Reserved Capacity
The report outlines four stages of AI infrastructure deployment.
Companies typically begin with serverless inference APIs because they can be deployed quickly and require little infrastructure management. As workloads grow, some move to on-demand GPUs for greater control. Predictable and sustained workloads can then shift to reserved infrastructure, followed by hybrid architectures that distribute workloads across multiple tiers.
Futurum says agentic AI could shorten this progression because higher token consumption can make the economics of production workloads change sooner. A workload that was economical on a token-metered API during experimentation may require a different infrastructure model once it operates continuously across an enterprise.
The report recommends considering reserved bare-metal infrastructure when utilization is sustained and predictable, particularly when baseline utilization is above roughly 60%, token demand rises from thousands to millions, specialized model-serving configurations are required, or data residency and privacy requirements limit hyperscaler options.
Bursty or unpredictable workloads, meanwhile, may remain better suited to on-demand infrastructure, particularly when they depend heavily on native hyperscaler services or are too small to justify the additional management associated with dedicated infrastructure.
AI-First Cloud Expands
Futurum forecasts that AI-first cloud infrastructure — including specialized AI clouds, bare-metal GPU providers and GPU marketplaces — will be the fastest-growing infrastructure tier in 2026, with spending projected to increase 107.1% during the year.
The category is forecast to grow from $45.2 billion in 2025 to $375.2 billion by 2030. Futurum expects it to reach 62% of the size of the hyperscaler tier by 2030.
The research identifies dedicated hardware, configuration flexibility, cost optimization and greater control over model-serving environments as factors supporting the expansion of this segment.
QumulusAI Positions Reserved Capacity
QumulusAI is positioning its reserved bare-metal GPU infrastructure as the production tier within a broader hybrid architecture. The company provides dedicated and virtualized GPU infrastructure for AI inference and training, with capacity based on reserved hardware rather than per-token charges.
“Per-token pricing makes sense when companies are experimenting, but the economics change quickly when AI applications move into sustained production,” said Michael Maniscalco, CEO of QumulusAI. He said enterprises can combine flexible services for experimentation with reserved capacity for predictable, high-volume workloads.
Futurum’s research includes examples from AI infrastructure and application companies that have adopted reserved capacity. Runpod, an AI developer cloud, uses reserved QumulusAI capacity, while Amberd.ai built a production offering around private, open-source large language models running on QumulusAI bare metal.
Infrastructure Economics Move Closer to the Application
For enterprises, the issue extends beyond the price of individual tokens. Inference workloads can run continuously, making utilization, latency, hardware availability and software optimization important components of the overall cost structure.
Futurum notes that inference differs from training because it is often sustained, distributed and latency-sensitive. Organizations may therefore need to optimize serving engines, batching, quantization, KV-cache management, observability and workload routing alongside the underlying GPU infrastructure.
“Enterprises are not choosing a single infrastructure model for AI,” said Brendan Burke, research director for semiconductors, supply chain and emerging technology at Futurum Research and author of the report. He said organizations are building hybrid environments that retain on-demand flexibility while moving predictable production workloads to reserved capacity.
The research was based on Futurum’s market forecasts and survey data from the first half of 2026, along with analyst-led interviews conducted in June with executives from Runpod, Qubrid AI and Amberd.ai. Futurum also discloses that it provides research, analysis, advisory and consulting services to technology companies, including companies mentioned in the report.






