HP Pushes AI Inference Beyond the Data Center with NVIDIA GB300 and Red Hat

ai factory nvidia gb300

HP wants to bring enterprise-grade artificial intelligence inference closer to where the data is generated. The company is working with Red Hat and NVIDIA on a platform that will combine HP ZGX Fury, the NVIDIA GB300 Grace Blackwell Ultra superchip, and Red Hat AI Factory to run models and agents locally, without every workload having to travel to a data center or a public cloud.

HP’s new AI platform in 30 seconds

  • HP is developing a platform with Red Hat and NVIDIA to run AI inference on-premises and at the edge.
  • HP ZGX Fury uses the NVIDIA GB300 Grace Blackwell Ultra and reaches up to 20 PFLOPS FP4.
  • The system has 748 GB of coherent memory shared between CPU and GPU.
  • Red Hat AI Factory will provide management, orchestration, and deployment of models and agents.
  • ZGX Fury is already orderable, but the full joint solution is still in development.

The announcement goes beyond a new workstation. HP is trying to move part of the architecture currently associated with AI servers into on-premises environments, combining it with tools to manage models, applications, and agents consistently across on-premises facilities and the cloud.

That can be especially interesting for companies that need to process information close to factories, stores, laboratories, offices, or sites with limited connectivity. It also applies to those that don’t want to send certain data to external services for privacy, regulatory, or sovereignty reasons.

The proposal isn’t meant to fully replace the data center. HP is presenting a hybrid model in which each workload can run wherever it makes the most sense.

A GB300 with 748 GB of coherent memory in a liquid-cooled tower

The most striking component is HP ZGX Fury, a system HP describes as an AI station, and one that technically sits quite far from a conventional workstation.

Its NVIDIA GB300 Grace Blackwell Ultra Superchip combines a 72-core Arm Neoverse V2 Grace CPU with a Blackwell Ultra GPU.

The specs published by HP list 496 GB of LPDDR5X memory for the CPU and 252 GB of HBM3e for the GPU, with the latter offering 7.1 TB/s of bandwidth.

In total, the architecture provides 748 GB of coherent memory.

SpecHP ZGX Fury
SuperchipNVIDIA GB300 Grace Blackwell Ultra
CPUGrace, 72-core Arm Neoverse V2
CPU memory496 GB LPDDR5X
GPU memory252 GB HBM3e
Total coherent memory748 GB
AI performanceUp to 20 PFLOPS FP4
Networking2 × QSFP112 at 400 Gbps
CoolingLiquid
Form factorTower, also rack-mountable
Initial OSUbuntu

HP says this capacity makes it possible to run models with up to a trillion parameters locally, though that claim depends on the model’s quantization and doesn’t mean any model of that size can run at any precision or configuration.

The company also places local fine-tuning around models in the 100-billion-parameter class.

According to HP’s specs, the system reaches up to 20 PFLOPS of FP4 performance, a precision format specifically aimed at accelerating AI workloads.

That said, theoretical PFLOPS figures don’t allow for a direct comparison of real-world performance across models, applications, and accelerators either. Inference speed will depend on factors such as the model used, quantization, context length, number of concurrent users, and the software serving it.

The machine isn’t especially small, either. The commercial configurations HP has published start at roughly 41 kilograms and use liquid cooling. Its form factor allows it to be used as a tower or mounted in a rack.

That combination helps clarify the market HP is targeting: teams that need capacity similar to an AI server but want to place it directly in an office, lab, department, or remote site.

Red Hat turns local hardware into a small AI platform

The second part of the deal tries to solve a less visible problem than buying a GPU: operating the infrastructure.

Having a GB300 on hand doesn’t by itself provide an enterprise platform for deploying models.

HP wants to pair ZGX Fury with Red Hat AI Factory with NVIDIA, built on Red Hat Enterprise Linux and Red Hat OpenShift and designed to manage AI models, agents, and applications across hybrid infrastructure.

The offering also incorporates technologies from NVIDIA AI Enterprise and Red Hat AI Enterprise.

The stated goal is to let different workloads share the same system while maintaining isolation, governance, and operational control. It also aims to improve GPU utilization through CUDA libraries, resource scheduling, and orchestration of multi-GPU workloads.

There’s an important distinction between the hardware and this joint platform.

HP ZGX Fury is already available to order and is certified for Red Hat Enterprise Linux, but the full solution built on Red Hat AI Factory is still upcoming.

HP will let customers evaluate it through isolated environments or sandboxes running on its devices before moving workloads to production. The company hasn’t yet announced dates, locations, supported configurations, or access terms for those trials.

So the announcement shouldn’t be read as though the entire joint platform were already commercially available.

Inference starts moving out of the big data centers

The proposal reflects a broader trend in AI infrastructure.

Training large models will keep requiring huge clusters of accelerators. Inference poses a different problem, because it happens continuously, every time a person, application, machine, or agent uses the model.

Sending all of those requests to a cloud region isn’t always the best solution.

A factory may want to analyze images from a production line in near real time. A hospital may handle information it prefers to keep within its own facilities. A government agency may need systems disconnected from the internet, and an engineering firm may want to give its developers access to internal models without paying for every token consumed externally.

Local processing can ease some of those problems, though it introduces others.

The company becomes responsible for acquiring and operating high-cost hardware, power consumption, cooling, storage, and upgrades. And physically keeping data on-site doesn’t automatically make a platform sovereign or guarantee regulatory compliance.

Software, identities, external connections, model governance, and data management all remain relevant.

HP specifically mentions manufacturing, engineering, software development, retail, healthcare, public administration, and distributed operations among the first planned use cases.

One of the clearest examples is industrial machine vision. Images from a production line can be analyzed locally to detect defects without continuously streaming large volumes of video to a remote cloud.

Another is agent-driven software development. A team could run models locally for coding, testing, and evaluation while keeping code and other corporate information within its own infrastructure.

The approach also partly changes the economics of AI consumption.

API-based services typically charge based on tokens processed. In-house infrastructure replaces part of that variable cost with investment and relatively more predictable operating expenses.

That doesn’t mean running AI locally is necessarily cheaper. It will depend on how much the system is used, its purchase price, electricity, maintenance, staffing, and useful life. An underused machine can end up far more expensive per token than buying external capacity.

For intensive, permanent, predictable workloads, the equation can look different.

That’s precisely the market HP is trying to cover with ZGX Fury: placing an amount of memory and compute power that until recently was associated with the data center directly next to the equipment that consumes AI, while Red Hat provides a common layer to manage it.

The cloud will remain part of that architecture. The difference is that the center of gravity for inference is starting to spread across the cloud, the data center, and much more powerful local systems.

FAQ

What is HP ZGX Fury?

HP ZGX Fury is an AI station for local development and inference based on the NVIDIA GB300 Grace Blackwell Ultra. It offers up to 748 GB of coherent memory and up to 20 PFLOPS of FP4 performance, according to HP.

Can it run trillion-parameter models?

HP says ZGX Fury can host models with up to a trillion parameters locally using FP4 quantization. Real-world feasibility and performance will depend on the specific model, its architecture, and configuration.

What does Red Hat AI Factory add?

It provides the enterprise software layer for deploying and managing AI models, agents, and applications in hybrid environments, built on technologies such as Red Hat Enterprise Linux and OpenShift.

Is the joint HP, Red Hat, and NVIDIA platform already available?

HP ZGX Fury is already available to order and is certified for Red Hat Enterprise Linux. The full integration planned with Red Hat AI Factory doesn’t yet have a firm public availability date.

For more on how NVIDIA’s Blackwell-generation superchips are reaching workstation-class hardware, see how MSI brought NVIDIA GB300 to the desktop with its XpertStation WS300.

Scroll to Top