Broadcom takes multi-tenant model sharing GA and turns the private cloud into a hyperscaler rival
On September 3, 2026, Broadcom updated VMware Cloud Foundation to 9.1.1 with multi-tenant model sharing in general availability, a gateway to more than 150 models, and a 2.6x jump in Kubernetes scale. For inference on regulated data, compare the total cost of a consolidated on-premises GPU estate against a hyperscaler’s per-token price before you commit.
August 31, 2026. Broadcom opens VMware Explore 2026 and unveils the VMware Private AI Cloud. September 3. VMware Cloud Foundation 9.1.1 reaches general availability. September 12. The first analyses put a number on the shift. Why it matters: for the first time, a virtualization vendor is openly arguing that a private cloud can rival AWS, Azure, and Google Cloud on AI workloads — not on raw power, but on the economics of GPUs.
A “point” release that changes the conversation
VCF 9.1.1 is technically a point release, four months after the general availability of VCF 9.1 in May 2026. But behind the version number, Broadcom has shipped capabilities that had been sitting in tech preview since spring — and they reposition the platform against the hyperscalers.
The most visible is multi-tenant model sharing moving to general availability. Before this release, an enterprise deploying a large language model on VCF often ended up provisioning one instance per business unit or project — meaning duplicated GPU capacity, duplicated licensing, and duplicated operational overhead. Multi-tenant sharing now lets a single deployed model serve multiple tenants, with data privacy and isolation enforced through Kubernetes namespaces. For a platform team serving five internal groups from the same base model, that is the difference between one GPU cluster and five.
GPU consolidation is the real economic lever
The reasoning reads directly in FinOps logic. On owned infrastructure rather than per-token billing, AI is not expensive because it is fast: it is expensive when GPUs sit idle. Multi-tenant sharing attacks exactly that waste. A pooled GPU estate only pays off if its utilization rate stays high, and namespace-level isolation is what makes sharing safe enough to attempt at scale.
It is a meaningfully different architecture from the hyperscalers. On Amazon Bedrock or Google Vertex AI, multi-tenancy is handled at the account or API-key level, not inside a shared cluster. Broadcom pushes isolation down into the cluster, which changes the compliance calculus for healthcare, banking, and government bodies restricted from sending data to a public cloud API.
The FinOps calculation you have to make
The decision does not hinge on a benchmark but on a utilization threshold. The reasoning is easy to lay out. A GPU server you bought has an amortized cost per hour, whether or not it is serving; a hyperscaler bills by consumption, per token or per GPU-hour. The tipping point is the utilization rate: below a certain threshold, on-demand rental is cheaper; above it, ownership wins.
Multi-tenant sharing moves that threshold. If one model serves five teams instead of five separate instances, the utilization rate climbs mechanically, and the amortized share of each token falls. The NVMe memory tiering announced at VMware Explore pushes the same way: by extending effective memory capacity without buying more RAM, it lowers the entry cost of an inference server. Together, these two levers bring the on-premises total cost of ownership closer to a hyperscaler’s per-token price — for continuous, predictable workloads.
The corollary is just as clear. An elastic, sporadic, or spiky workload remains structurally favorable to the public cloud: ownership only justifies itself if you amortize the hardware over a continuous stream of use. A purchased GPU cluster that sleeps for half the night is more expensive than any per-token rate.
The numbers Broadcom is pushing
The most-quoted figure of the cycle is a 2.6x jump in Kubernetes cluster scale over earlier preview builds, alongside a 75% reduction in deployment time and a 75% shorter upgrade window, according to NetworkWorld’s review. Those gains come from improvements to the VMware Kubernetes Service and virtualized load balancing through Avi Load Balancer paired with vDefend, which remove dedicated hardware appliances in front of inference endpoints.
The second number is about maintenance. Live patching now covers roughly 80% of common upgrade scenarios without a maintenance window. For a team running production inference clusters, where every scheduled outage collides with GPU-utilization targets, that is the detail that decides how often patches actually get applied.
A gateway to 150+ models, still in preview
Alongside multi-tenant sharing, Broadcom previewed an AI gateway capable of brokering access to more than 150 open-source and open-weight models, with per-application authorization controls. The underlying thesis is clear: enterprises do not want to commit to a single model vendor, but to a governed catalog where they switch between models on cost, latency, or compliance.
The difference from the catalogs of Amazon, Microsoft, and Google comes down to where execution happens: Broadcom’s version runs entirely inside the customer’s own infrastructure. For data that must not leave the perimeter, that is the distinction that matters — even though the feature remains in preview with no announced GA date, as do GitOps support and native object storage.
What is still missing, and the lock-in risk
The picture has its shadows, and they should be named. GitOps and native object storage remain in tech preview: two building blocks that platform engineering teams now treat as prerequisites, not options. And while the model gateway promises not to lock you into a model vendor, it locks you into Broadcom: the whole stack — hypervisor, Kubernetes, load balancing, AI gateway — is sold by one vendor under a VCF subscription that many long-time customers already consider pricier than their old vSphere licenses.
This is the real counter-argument. GPU economics are attractive, but they come at the price of reinforced dependence on a vendor that has already shown it knows how to raise prices. An honest FinOps calculation must therefore include not just the cost of silicon, but the cost of the license and of the exit — the cost of migration if you ever decide to move back.
The pressure of an end-of-support calendar
This push into AI does not happen in a vacuum. TechTarget frames Broadcom’s position bluntly: the vendor is racing the clock on end-of-support timelines for its legacy VMware estate, and is using the new AI and Kubernetes capabilities as the incentive to move. Since finalizing the roughly $69 billion acquisition of VMware in November 2023, Broadcom has restructured licensing around subscriptions rather than perpetual vSphere licenses — a shift that angered many long-time customers pushed toward costlier VCF bundles.
The competitive context frames the whole thing. Amazon, for its part, expanded Aurora Serverless v4 performance gains to five new regions on August 31 — the same race toward “AI-native” infrastructure, but through the public cloud and consumption model rather than the private cloud and ownership model.
Verdict
If your data is regulated and your GPU estate is genuinely utilized, VCF 9.1.1’s pitch deserves to be priced out: a consolidated cluster with multi-tenant sharing, live patching, and virtualized load balancing can make on-premises inference competitive against a hyperscaler’s per-token price — provided utilization stays high, because an owned infrastructure that sleeps is the most expensive of all. If your workload is elastic or sporadic, consumption billing on AWS, Azure, or GCP remains unbeatable: ownership only justifies itself through continuous use. If you are still on VCF 8 or earlier, the question is not whether you migrate but when — and the end-of-support window makes that migration your most urgent cost decision, well ahead of which AI model you pick.
References
- Broadcom — VMware Cloud Foundation brings leading AI models to the Private AI Cloud, August 31, 2026
- Shattered.io — Broadcom VCF 9.1.1 Rivals AWS AI Cloud at 2.6x Scale, September 12, 2026
- TechTimes — VMware Explore 2026: Broadcom Solves AI Server DRAM Crisis With NVMe Memory Tiering, September 1, 2026
- Virtualization Review — VMware Explore 2026: Highlighted Announcements, August 31, 2026
- ETTAYEB — Cloud FinOps: taking back control of the bill