Skip to content
Breaking:

Agentic AI Token Consumption Pushes Enterprises Toward Reserved Compute Infrastructure

A study sponsored by QumulusAI shows autonomous AI workflows consume up to 100 times more tokens per task than standard inference calls, forcing a reassessment of on-demand cloud pricing.

By The Company Wire4 min read
Share
QumulusAI — Agentic AI Token Consumption Pushes Enterprises Toward Reserved Compute Infrastructure
QumulusAI — Agentic AI Token Consumption Pushes Enterprises Toward Reserved Compute Infrastructure. Photo: SiliconANGLE.

Autonomous artificial intelligence systems are generating vastly larger volumes of tokens than traditional querying tools, challenging the viability of pay-as-you-go pricing for production workloads. According to a research report titled “The Off Ramp From Per-Token Pricing” published by Futurum, sponsored by neocloud provider QumulusAI Inc. and reported by SiliconANGLE , agentic AI workflows consume between 10 and 100 times more tokens per task than simple inference calls.

Per-token application programming interface models allow software teams to prototype rapidly without upfront procurement cycles, but linear consumption costs scale quickly when systems reach production scale. While standard chatbots only generate answers to direct prompts, agentic workflows execute multi-step planning, tool invocations, intermediate verification, retries, handoffs, and iterative summaries. Futurum forecasts that agent and reasoning inference workloads will grow 219% this year, with total global enterprise inference spending projected to climb from $120 billion in 2025 to $885 billion by 2030.

This unpredictable spending pattern is already straining corporate IT budgets. Mazda Marvasti, co-founder and chief executive of Amberd.ai, noted in the report that variable usage rates cause operational costs to escalate rapidly across broad enterprise deployments. Marvasti indicated that several organizations have abandoned internal automation tools simply because finance and IT leaders could neither forecast nor justify the ongoing variable expenses.

To mitigate variable consumption costs, enterprises are shifting steady workloads away from pure on-demand public cloud instances. In a Futurum survey of 824 AI decision-makers, reserved and owned infrastructure accounted for 66% of total AI compute consumption, compared with 19% on on-demand cloud services. Furthermore, 59% of respondents reported running primary AI workloads outside hyperscaler public clouds, utilizing colocation facilities, dedicated bare-metal providers, or private data centers.

Some software providers are utilizing custom virtualization architectures on reserved hardware to maintain margin control. Amberd.ai, for example, partitions an eight-GPU Nvidia H200 server hosted on QumulusAI bare metal into four virtual environments of two GPUs each, tiering client workloads by latency tolerance. Marvasti stated that two customers cover the cost of an entire server, enabling the company to host up to 35 customers on a single unit and turn subsequent capacity into profit.

Futurum recommends reserved bare metal for sustained enterprise workloads that maintain predictable utilization above roughly 60%. However, the report cautions that running private infrastructure requires specialized engineering—including serving engines, batching, quantization, and key-value cache management—without which organizations may encounter separate operational overheads. The strategy also depends on workload architecture: while open-weight models can run on private bare metal, proprietary frontier models remain constrained to vendor APIs and hyperscaler pricing structures.

Sources

  1. SiliconANGLE

Company: QumulusAI

Written by

The Company Wire

Newsroom · San Francisco

Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.