A host of cost items and compliance issues can be eliminated in a single step. That’s the promise behind VMware Private AI Cloud, a solution from Broadcom that combines virtualization and LLMs to provide a more secure alternative for AI usage.
Private AI Cloud is inextricably linked to VMware Cloud Foundation 9, the infrastructure layer of VMware by Broadcom that integrates all the underlying IT and AI systems. This is now more true than ever, with validated LLMs and “tokenomics”-focused configurations designed to keep AI usage affordable.
Broadcom’s approach is simple. It boils down to this: bring the model to the data, not the other way around. With VMware Private AI Cloud, the company aims to offer organizations a path to build AI workloads within a mature and compliant IT system, where the workloads are run and managed entirely within the organization’s own infrastructure when necessary. Customers choose their own hardware, accelerators, and models.
Model to the data, not data to the model
This goes deeper than just a basic alternative to API usage. Last year, Broadcom made VMware “AI-native” with VCF 9 and Private AI Services, making features such as GPU monitoring, model runtime, and vector databases standard parts of the platform. The new announcements build directly on this foundation and capitalize on the fact that open AI models are now much more competitive and usable than they were when VCF 9 launched in the middle of last year.
VMware AI Factory is the software-defined foundation of the Private AI Cloud, and for this, a partnership with AMD was specifically chosen. Broadcom claims that the time between deploying a physical server and deploying the first AI model has been reduced from weeks to hours. Hardware provisioning, software stack deployment, and lifecycle management are fully automated. Although AMD’s name appears eight times in the announcement, we should not forget that all of this is also possible with Nvidia, with whom a partnership was already in place.
Many features
To further accelerate AI deployment, Broadcom is partnering with MetalSoft for heterogeneous bare-metal automation. Administrators can provision or reconfigure physical servers from multiple vendors directly through the VCF management console.
In addition, Broadcom is introducing private AI services such as multi-tenant model sharing via isolated namespaces, an AI Gateway with prompt routing and token rate limiting, and secure AI sandboxes for code generated by agents. GPU resources are pooled and shared, allowing multiple models to run on the same hardware. The logical partitioning of GPUs has often been a pain point, forcing organizations to provision per GPU rather than pooling all computing power into a single stack. The latter is now possible.
Over 150 models validated
VCF customers can now run more than 150 open-source and commercial models, with vLLM as the default runtime. Broadcom specifically validated Nvidia’s Nemotron 3, Google DeepMind’s Gemma 4, NEC’s Japanese model cotomi, Alibaba’s Qwen3.8-27B, and Z.ai’s GLM 5.2. Support for AMD, Intel, and Nvidia hardware is intended to limit vendor lock-in.
For the agentic layer, Broadcom unveiled AgentMinder, a central management layer that treats AI agents as enterprise identities. Their permissions are tied to a specific task, approved tools, and authorized resources, with runtime policy enforcement and audit logging. The Tanzu Platform employs a deny-by-default runtime: agents are not granted access to APIs, networks, MCP servers, or the internet unless explicitly permitted.
On the security front, vDefend extends zero-trust-based lateral security to agentic workloads, including detection of “shadow AI,” that is, the unauthorized or unknown use of AI by employees. The Avi Load Balancer is designed to prevent agents from accessing unauthorized tools or exfiltrating sensitive data.
According to Broadcom, there is no doubt that customers want these features. Its own “Private Cloud Outlook 2026” report shows that 56 percent of the organizations surveyed are already running AI inference on a private cloud basis or plan to do so. For those same workloads, public cloud usage dropped to 41 percent over the course of a year, a striking sign of skepticism regarding the AI approach pursued by the hyperscalers.