VMware Brings Nemotron, Gemma, Qwen, and GLM to Its Private AI Cloud

Broadcom has expanded the artificial intelligence capabilities of VMware Cloud Foundation (VCF) with the validation of several models from NVIDIA, Google DeepMind, NEC, Alibaba Cloud, and Z.ai to run on private infrastructure. The announcement, made at VMware Explore 2026, aims to let companies offer inference and models as a service from their own data centers or private clouds, keeping greater control over data, infrastructure, and costs. It follows Broadcom’s related unveiling of VMware AI Factory, a broader automation layer built on the same VCF platform.

Private AI on VMware in 20 seconds

  • VCF validates models from NVIDIA, Google DeepMind, NEC, Alibaba Cloud, and Z.ai.
  • Among them are Nemotron 3, Gemma 4, cotomi, Qwen3.8-27B, and GLM 5.2.
  • vLLM also lets customers run more than 150 open models, according to Broadcom.
  • VCF supports infrastructure based on AMD, Intel, and NVIDIA.
  • The goal is to bring inference and AI agents into private environments.

The news reflects a trend that’s starting to define the next phase of enterprise AI. After experimenting mainly through APIs and public services, some organizations are now weighing which workloads are worth running on infrastructure they control.

Broadcom states in its Private Cloud Outlook 2026 that 56% of surveyed companies are already running or plan to run AI inference in production on private cloud. That figure comes from the company’s own study and should be read within that context, but it helps explain the direction VMware Cloud Foundation is taking.

VCF’s goal is for traditional virtual machines, Kubernetes, agentic applications, and model inference to share a single infrastructure platform.

Five model families validated for VCF

The most concrete piece of news is the validation of several models to run on VMware Cloud Foundation.

Broadcom has announced validated compatibility with NVIDIA Nemotron 3, Google DeepMind Gemma 4, NEC cotomi, Alibaba’s Qwen3.8-27B, and Z.ai’s GLM 5.2.

Not all of them address the same kind of need.

Nemotron 3 is aimed especially at agentic applications and combines, according to Broadcom’s description, a hybrid Mamba-Transformer architecture with Mixture of Experts (MoE), multimodal capabilities, and a context window of up to one million tokens.

Gemma 4 represents Google DeepMind’s family of open models within the platform. Broadcom highlights its multimodal orientation and its use for building applications and agents that can run locally.

The presence of cotomi, developed by NEC, introduces a different case. It’s a model built especially for Japanese and enterprise scenarios in that market. NEC claims it achieves a 40% improvement in token efficiency, a figure attributed to the vendor.

Alibaba contributes Qwen3.8-27B, described in the announcement as a dense multimodal model with 27 billion parameters aimed at coding, research, professional tasks, and long-running agents.

Finally, GLM 5.2 from Z.ai — formerly known as Zhipu AI — is aimed at reasoning, coding, and multi-step autonomous processes.

ModelProviderFocus highlighted by Broadcom
Nemotron 3NVIDIAAgents, multimodality, and long contexts
Gemma 4Google DeepMindLocal AI, multimodality, and agents
cotomiNECJapanese and enterprise applications
Qwen3.8-27BAlibaba Cloud / QwenCode, vision, research, and agents
GLM 5.2Z.aiReasoning, code, and autonomous workflows

Validation doesn’t mean these models need VMware to run. They can be deployed through other infrastructure and tools. What Broadcom offers is a tested environment to integrate them into VCF and manage them alongside the rest of an enterprise’s workloads.

vLLM extends the catalog well beyond five models

The announced list isn’t the limit of models that can be used, either.

VMware Cloud Foundation adopts vLLM as its default runtime for serving models, which lets Broadcom claim that customers can run more than 150 compatible open models.

vLLM has become one of the most widely used projects for serving large language models because it incorporates techniques designed to efficiently manage memory, concurrency, and token generation.

Its presence matters to keep the infrastructure from being tied exclusively to a closed catalog.

A company could use a model Broadcom specifically validated for certain applications and other vLLM-compatible models for different projects.

The choice will also depend on the hardware available.

Broadcom points out that VCF supports mixed compute configurations with AMD, Intel, and NVIDIA, so the strategy isn’t limited to a single type of processor for every workload.

That can matter in data centers where there’s already significant investment in infrastructure and companies want to add AI gradually, rather than build a completely separate platform.

Private inference is gaining ground over training

The announcement is aimed specifically at inference, not at competing with the huge infrastructure used to train foundation models from scratch.

That distinction matters.

Training a frontier model can require thousands or tens of thousands of accelerators over long periods — a scale that’s out of reach for the vast majority of companies.

Running an already-trained model to handle internal requests has a different economics.

An organization can use it for corporate assistants, document search, code generation, information analysis, or agents connected to enterprise applications.

When query volume is stable enough, deploying that inference on owned infrastructure can become an option worth weighing against ongoing per-token payments to external providers.

That doesn’t mean private cloud is automatically cheaper.

You have to factor in GPUs, servers, power, cooling, storage, licensing, staff, effective hardware utilization, and technology refresh cycles. An underused GPU can make consuming an external API considerably cheaper.

That’s why Broadcom introduces the concept of AI tokenomics: analyzing the real cost of producing and consuming tokens depending on where the model runs.

The decision is starting to resemble other long-standing infrastructure debates. Some variable workloads fit well on public services; others, stable and predictable, can justify dedicated capacity.

Data is the other argument for bringing AI in-house

Cost is only part of the decision.

Privacy and data governance explain much of the interest in private AI.

An enterprise agent becomes more useful when it can query internal documentation, knowledge bases, applications, code, or customer information — precisely the kind of data an organization can find harder to send outside certain perimeters.

Running the model within controlled infrastructure allows for architectures where a larger share of processing stays inside the private environment.

That doesn’t automatically make a local deployment secure or compliant.

Security will depend on identities, permissions, segmentation, applications, model configuration, and the systems it can access. It will also be necessary to control what information ends up in prompts, logs, and vector databases.

But having the model within the same environment gives the organization more options for defining where that data flows.

Broadcom ties this feature to sovereignty and governance, especially in sectors subject to regulatory constraints.

VMware wants to offer models as an internal service

Another element of the announcement is the concept of Model as a Service.

The idea is to avoid having every department deploy and manage its own inference servers.

The infrastructure team can maintain different models on VCF and offer them internally as services to developers and departments.

That way, one team could consume Gemma for an application, another could use Nemotron for agents, and a third could deploy a specialized model, while the infrastructure remains centrally managed.

It’s an evolution similar to what Kubernetes brought to enterprise applications: the platform tries to abstract away part of the hardware complexity and provide resources on demand to those building services.

AI adds a further difficulty: GPUs are expensive and limited resources, so improving their utilization can have a direct effect on cost.

Broadcom also points to results under MLPerf Inference v5.1 to argue that VCF can deliver performance comparable to bare metal in the evaluated scenarios. That reference should be understood within the specific configurations submitted to the benchmark, not as a guarantee that any virtualized workload will perform exactly like a direct install on hardware.

Private cloud is trying to become an AI platform too

Broadcom’s strategy with VCF is starting to go beyond maintaining existing virtualized applications.

The company is trying to turn VMware Cloud Foundation into a platform where traditional infrastructure and new AI workloads coexist.

For many companies, that coexistence can be more practical than building a separate environment exclusively for artificial intelligence.

Virtual machines aren’t going away because an organization adds agents. Neither are the databases, Java applications, Windows systems, or Kubernetes platforms it already uses.

The question is whether existing infrastructure can also absorb models, GPUs, and new inference services without creating another technology island.

Validating Nemotron, Gemma, cotomi, Qwen, and GLM aims to answer exactly that need.

The diversity of providers also matters. Broadcom isn’t proposing a private cloud tied to a single model. Its pitch is to provide a layer where the company can switch models as the market evolves.

In an industry where the differences between models can narrow quickly and new versions appear every few months, avoiding excessive dependence on one specific model can end up being just as important as picking the best one available at a given moment.

VMware Cloud Foundation is trying to occupy that middle layer: below sit servers, CPUs, and GPUs; above sit models and applications; in between sits a private platform in charge of providing the resources, Kubernetes, virtualization, and services needed to operate them.

The VMware Explore 2026 announcement shows just how far AI is moving into the traditional enterprise infrastructure market. The next debate no longer revolves only around which model to use, but also where to run it, how much each token costs, and what data is allowed to leave the organization’s infrastructure.

Frequently Asked Questions

What AI models has Broadcom validated for VMware Cloud Foundation?

Broadcom has announced the validation of NVIDIA Nemotron 3, Google DeepMind Gemma 4, NEC cotomi, Alibaba’s Qwen3.8-27B, and Z.ai’s GLM 5.2.

Is VCF limited to those five models?

No. Broadcom uses vLLM as its default runtime and states that VMware Cloud Foundation can run more than 150 compatible open models.

What’s the point of running an AI model on a private cloud?

It allows for greater control over infrastructure and data flow and can be attractive for stable inference workloads. Whether it makes economic sense depends on hardware cost, utilization, power, operations, and the external services available as alternatives.

Does VMware Cloud Foundation need NVIDIA GPUs exclusively?

No. Broadcom says VCF supports compute configurations with AMD, Intel, and NVIDIA technologies, although actual compatibility and performance will depend on each model and configuration.

Scroll to Top