Enterprises Pivot to Private AI Clouds as Token Costs and Data Sovereignty Concerns Mount
Broadcom executive Chris Wolf outlines why IT operations teams are moving production inference workloads back to owned infrastructure using software-driven architectures.

As enterprise artificial intelligence initiatives transition from early experimentation to full-scale production inference, corporate technology departments are encountering significant operational challenges in public cloud environments. Questions surrounding data governance, regulatory compliance, regional digital sovereignty, and rising token costs are prompting organizations to bring core intelligence capabilities back onto on-premises infrastructure. This structural shift is turning traditional private clouds into dedicated private AI platforms designed specifically to meet the stringent performance and control requirements of live production workloads.
This operational migration places new burdens on enterprise information technology operations teams, which must maintain a complex blend of artificial intelligence systems on shared physical hardware. Data center managers are required to support large frontier models alongside compact small language models (SLMs) and autonomous agent swarms without driving up physical server counts, energy consumption, or software licensing fees. In an interview broadcast on SiliconANGLE's livestreaming studio theCUBE during the VMware Explore 2026 conference, Chris Wolf, global head of AI and advanced services for the VMware Cloud Foundation Division at Broadcom Inc., highlighted this resource consolidation as a primary architectural challenge facing modern enterprises.
Wolf emphasized that organizations must adopt a pragmatic, multi-tiered approach to model deployment rather than relying exclusively on cloud-hosted frontier models. "You have sovereignty considerations. You have tokenomics considerations as well. This doesn’t mean don’t use frontier models. It means be practical," Wolf told theCUBE host John Furrier, in reporting first published by SiliconANGLE. "Use frontier models where it makes sense, where [you] need deep reasoning. Use specialized models, local SLMs, where they make sense as well. You’re really seeing this breadth of coverage happening in the industry — and now IT operations is caught in the middle of all of this."
The rapid adoption of agentic AI workloads has further intensified infrastructure demands, requiring dedicated virtualized environments capable of supporting dynamic multi-agent deployments securely. Enterprise setups now require warm pools of isolated virtual machines that can instantly instantiate autonomous software agents while enforcing strict boundaries to prevent privilege escalations or system escapes. At the same time, managing graphics processing unit (GPU) memory has emerged as a central bottleneck, requiring precise key-value cache placement across accelerators alongside multi-tier storage caching. Broadcom is positioning its VMware Cloud Foundation platform as a central abstraction and pooling layer to manage these complex hardware constraints.
A frequent mistake made by enterprise buyers is prioritizing hardware procurement over software integration, according to Wolf. Many IT organizations purchase physical servers before selecting their underlying software stack, creating major compatibility and operational flexibility issues. "‘Buy your hardware first, figure out the software later’ — no, that’s a horrible idea, because you have to make sure that your software choices are compatible with the hardware you bought," Wolf explained during the broadcast. "Software is what’s giving you the ability to have autonomy in terms of the accelerators you use, to have the flexibility to ensure that I can use cloud models when I need to [or] use local models when I need to."
This disconnect often leads to failed deployments and unexpected operational costs for organizations attempting to roll out AI infrastructure. "People were running into buyer’s remorse," Wolf noted, describing enterprise customers who transitioned from earlier cloud deployments. "They bought what they thought was this full turnkey solution, and as it turns out, it wasn’t."
To help enterprises avoid these integration hurdles, Broadcom is advocating for an "AI factory" operational model that automates infrastructure deployment end to end. The framework handles provisioning from raw bare-metal servers through model runtime layers, ultimately generating a declarative YAML file that documents the entire environment. Enterprise operators can then export this YAML file to replicate and clone identical compute clusters across distinct data centers, streamlining scaling operations without requiring manual reconfiguration.
Digital sovereignty concerns are also compelling enterprises and governments across North America, Europe, and Asia to demand strict control over their data planes, control planes, encryption keys, and hosted models. As regulatory environments become more complex and technology continues to evolve rapidly, Wolf advised IT executives to plan their infrastructure strategy for the next 18 months around maximum flexibility and software priority. "More than ever, they have to architect for the expectation of change. They can’t architect based on what looks good today, because the space is moving too fast," Wolf said. "Make sure software is at the forefront of your architecture and decision-making, and then go from there."
Sources
Written by
The Company Wire
Inside the companies building what’s next. Reporting on startups, technology, funding and the people shaping them.


