OpenNebula Validates Its Virtualization Platform for NVIDIA GB200 NVL72

OpenNebula has completed validation of its platform under NVIDIA’s hypervisor program for GB200 NVL72 systems, a step meant to show that AI workloads can run inside virtual machines while keeping direct access to the GPUs. The move addresses a question that keeps coming up in AI infrastructure: how to keep the operational advantages of cloud without being forced to run accelerators exclusively on bare metal.

OpenNebula and NVIDIA GB200 NVL72 in 20 seconds

  • OpenNebula has validated its virtualization platform for NVIDIA GB200 NVL72 systems.
  • The architecture uses PCI passthrough to assign GPUs and other devices directly to virtual machines.
  • The goal is to combine isolation, automation, and multi-tenancy with direct access to accelerator hardware.
  • OpenNebula says it’s currently the only European cloud platform validated under this NVIDIA program.
  • The integration is part of its pitch for AI Factories, HPC, and GPU infrastructure providers.

The validation is notable because it arrives as GPUs are reshaping some traditional architecture decisions. For years, high-performance computing (HPC) workloads, and later AI workloads, have favored bare metal to avoid any layer that might cut into performance or add latency.

The problem is that skipping virtualization also means giving up tools that have made cloud infrastructure easier to operate for years: tenant isolation, automated provisioning, resource policies, centralized management, and independent lifecycles for each workload.

Hardware evolution and technologies like PCI passthrough are narrowing that trade-off.

A virtual machine with direct GPU access

The most important technical piece of the approach is PCI passthrough.

In conventional virtualization, certain devices can end up behind layers of emulation or abstraction. With passthrough, a physical PCIe device can be assigned directly to a virtual machine so the guest operating system interacts with it directly.

Applied to a GPU, this means the VM still provides the isolation and management boundary, while the accelerator is handed directly to that workload.

That’s especially relevant for AI and HPC, where applications depend on high data throughput and efficient communication between CPU, GPU, storage, and network.

OpenNebula is thus aiming to keep a path to the accelerator close to what a bare-metal deployment would offer, without abandoning the operational model of a virtualized infrastructure.

For a cloud provider or a company managing hundreds of GPUs, that difference can matter a lot.

The question is no longer just about squeezing maximum performance out of each accelerator. It also means allocating resources across teams, isolating tenants, automating deployments, enforcing quotas, and tearing down environments once they’re no longer needed.

That’s where virtualization becomes relevant again, even in infrastructure built around GPUs.

NVIDIA GB200 NVL72 takes the problem to another scale

The validation was carried out on NVIDIA GB200 NVL72, an architecture designed for large-scale AI workloads.

These systems pack multiple Blackwell accelerators and Grace processors into an architecture connected via NVLink. The idea is that a large number of accelerators can work as a single large compute domain for training and inference of large models.

Managing this kind of infrastructure as a collection of traditional servers becomes less practical the moment it has to be shared across different projects or customers.

That’s why a new software layer is emerging around so-called AI Factories, designed to manage compute, GPUs, networking, and storage as pooled resources.

OpenNebula positions its platform squarely in that layer.

The company describes its offering as an open, vendor-neutral platform for building multi-tenant AI Factories, GPU services for neoclouds, HPC centers, and telecom operators.

Its integration with NVIDIA goes beyond the GB200 NVL72, too. OpenNebula documents compatibility or integration with technologies such as NVLink and NVSwitch, BlueField data processing units (DPUs), Spectrum-X Ethernet, and InfiniBand. It also integrates with NVIDIA Run:ai for scheduling and allocating GPU resources on Kubernetes.

Virtualization and bare metal are starting to coexist in AI Factories

The validation also reflects how AI infrastructure is changing.

In the early large GPU clusters, squeezing out maximum performance could justify fairly rigid setups. The server was dedicated to one workload, and software was installed directly on the operating system.

That model still makes sense in certain scenarios, especially when an organization needs to dedicate an entire machine to a single training run for long stretches of time.

But commercial infrastructure has other needs.

A GPU cloud provider might serve several different companies at once. A research center might split its accelerators across numerous projects. An organization might need separate environments for training, inference, development, and experimentation.

Using virtual machines makes it possible to create those boundaries without necessarily dedicating a separate management infrastructure to every single user. That’s the same trade-off explored in choosing between bare metal and virtualization for different workloads.

OpenNebula also adds other options for splitting up accelerators. Its platform supports NVIDIA MIG (Multi-Instance GPU) and vGPU technologies, designed to partition certain GPU resources when dedicating a whole accelerator to a single workload isn’t necessary.

That makes it possible to tell several models apart.

The most demanding workloads can get entire GPUs via passthrough. Other services can share accelerators using partitioning mechanisms. And nodes can still be part of infrastructure managed from a common cloud layer.

That flexibility can be especially useful for improving GPU utilization, one of the most expensive assets in new AI platforms.

A European technology inside NVIDIA’s program

OpenNebula also adds a European angle to the story.

The company says it’s currently the only European cloud platform validated under the NVIDIA AI Cloud-Ready ISV Program. Its documentation explicitly identifies OpenNebula’s full cloud infrastructure stack as validated on NVIDIA GB200 NVL72.

The precision matters: the claim refers to this specific NVIDIA validation category and program — it doesn’t mean OpenNebula is the only European technology compatible with NVIDIA GPUs, or the only one used in AI infrastructure.

For Europe, having its own management layers is also gaining relevance around sovereign AI and AI Factory projects.

The hardware may come from NVIDIA and other international manufacturers, but the layer that manages users, virtual machines, Kubernetes, networking, storage, and policies also determines who operationally controls the infrastructure.

OpenNebula is trying to occupy that space with an open platform usable in private infrastructure, HPC centers, operators, and cloud providers.

The company itself talks about avoiding technological dependency and keeping control over infrastructure and software. Those are OpenNebula’s commercial goals, though they line up with the broader European debate over technological sovereignty and the capacity to operate AI infrastructure independently.

The GB200 NVL72 validation now adds a further technical proof point for one of the most demanding scenarios OpenNebula is trying to cover.

It also points to a broader trend: bare metal and virtualization are increasingly not being treated as mutually exclusive options in AI.

When a GPU can be handed directly to a VM, the discussion shifts to how much isolation and automation each workload needs, which resources can be shared, and what real-world impact each architecture actually introduces.

In AI Factories that can bring together thousands of accelerators and numerous users, managing those GPUs well can end up mattering almost as much as having them in the first place.

Frequently asked questions

What has NVIDIA validated about OpenNebula?

OpenNebula has completed validation of its platform for NVIDIA-accelerated infrastructure, including its deployment on NVIDIA GB200 NVL72 systems under NVIDIA’s corresponding program.

How can a virtual machine access a GPU directly?

Through PCI passthrough, the PCIe device can be assigned directly to a virtual machine. The guest operating system gets hardware access while the VM keeps functioning as the isolation and management unit.

Does OpenNebula force all AI workloads to be virtualized?

No. The platform is designed to manage different infrastructure models, including virtualization, bare metal, and Kubernetes, plus technologies for partitioning or sharing GPUs.

Is OpenNebula the only European platform certified by NVIDIA?

OpenNebula says it’s currently the only European cloud platform validated through the NVIDIA AI Cloud-Ready ISV Program. That claim should be understood within the specific scope of that program, not as general exclusivity over NVIDIA technologies in Europe.

Scroll to Top