d-Matrix will integrate its next-generation Raptor XPU AI inference accelerators into NVIDIA’s MGX rack architecture via NVLink Fusion, as part of a multi-year collaboration aimed at bringing these specialized inference chips into large data centers without building an infrastructure stack from scratch. The first integrated systems are expected in the fourth quarter of 2027, while Raptor still has to clear an essential milestone: the company expects to finalize its design and send it to fabrication before the end of 2026.
The d-Matrix–NVIDIA alliance in 30 seconds
- d-Matrix will design Raptor to integrate directly with NVIDIA NVLink Fusion and MGX racks.
- Systems will combine technologies such as Vera CPU, BlueField-4, ConnectX-9, and Spectrum-X.
- Raptor is built for inference and uses an architecture that stacks DRAM on top of an SRAM compute chip.
- d-Matrix wants to pair its XPUs with NVIDIA Vera Rubin systems to split up different phases of inference.
- Initial availability is targeted for the fourth quarter of 2027.
The announcement is a window into a broader shift in the AI hardware market. NVIDIA is using NVLink Fusion to extend part of its infrastructure to processors built by other companies, including specialized accelerators, or XPUs, and custom CPUs.
For companies like d-Matrix, this solves a problem that goes well beyond building a good chip.
An accelerator meant for data centers needs high-speed connectivity, servers, racks, power delivery, cooling, node-to-node networking, and a supply chain capable of producing the whole system at volume. Designing each of those pieces independently can add years of work and raise the cost of getting into large-scale deployments.
d-Matrix wants to skip most of that path by adopting an architecture already built around NVIDIA.
Raptor Will Connect via NVLink and Use the MGX Ecosystem
The first result of the collaboration will be a rack system based on NVIDIA MGX, into which d-Matrix’s Raptor XPUs will be installed.
The company plans to incorporate NVIDIA components such as Vera processors, NVLink switches, BlueField-4 data processing units (DPUs), ConnectX-9 SuperNICs, and the Spectrum-X Ethernet network.
It’s also working with Astera Labs, a member of the NVLink Fusion ecosystem, to develop connectivity solutions designed to sustain high data throughput inside the system.
The approach splits communication into two tiers.
NVLink will be used for so-called scale-up — that is, for connecting accelerators within a high-speed, low-latency compute domain. Spectrum-X handles scale-out, linking different systems and racks across a larger facility.
| Component | Role in the system |
|---|---|
| d-Matrix Raptor XPU | Specialized inference |
| NVIDIA Vera CPU | General processing and coordination |
| NVIDIA NVLink Fusion | High-speed interconnect between accelerators |
| NVIDIA BlueField-4 | Infrastructure and network processing |
| NVIDIA ConnectX-9 | High-performance connectivity |
| NVIDIA Spectrum-X | Ethernet network for scaling across systems |
| NVIDIA MGX | Modular server and rack architecture |
| Astera Labs | Connectivity solutions for the integration |
The advantage d-Matrix is after is being able to focus more resources on the processor itself while using a rack, cooling, and connectivity infrastructure that already has a mature supply chain behind it.
That doesn’t turn Raptor into an NVIDIA GPU, nor does it mean NVIDIA will manufacture the accelerator. NVLink Fusion functions here as an integration point between third-party silicon and part of NVIDIA’s infrastructure platform.
A Memory Architecture Built Specifically for Inference
Raptor will succeed Corsair, d-Matrix’s inference platform that’s already in production.
The most interesting technical difference is in memory.
Generative models need to constantly move huge amounts of parameters and data between memory and compute units. That movement can become one of the main limits on inference speed, power consumption, and cost.
d-Matrix has spent years working on an architecture built precisely to shrink that distance.
Raptor will use a 3D stacking approach that places a DRAM memory chip next to an SRAM-based compute chip within a single vertical package, an approach the company graphically describes as a two-story structure.
The technical details of this technology were presented at Hot Chips 2026 by Sudeep Bhoja, co-founder and CTO of d-Matrix, and later published by IEEE.
The company says Raptor has more than 100 associated patents and states that the processor is being evaluated by major cloud providers and AI labs. It hasn’t publicly identified which companies are taking part in those evaluations in this announcement.
It’s also not a commercial product yet.
d-Matrix expects to tape out Raptor before the end of 2026. Tape-out marks the point where the design is finalized and sent to fabrication, but first silicon, testing, fixes, validation, and volume production still lie ahead after that.
That’s why the company is placing the first availability of Raptor systems integrated into MGX in the fourth quarter of 2027.
GPUs and XPUs Could Split a Single Inference Workload
The collaboration doesn’t necessarily position Raptor as a GPU replacement.
d-Matrix proposes using heterogeneous architectures in which different processors run the parts of a workload they’re best suited for.
One example is disaggregated inference for language models.
When someone sends a request to a model, there are two phases with different characteristics. During prefill, the system processes the full initial context. Then decode begins, during which it generates new tokens progressively.
Prefill needs a lot of parallel compute capacity. Decode, especially in interactive applications, is very sensitive to latency and to fast memory access.
d-Matrix proposes that NVIDIA Vera Rubin systems handle prefill while Raptor XPUs process certain decode workloads.
The company cites coding assistants, real-time chatbots, and voice agents as examples, where cutting the time between tokens can directly shape the user experience.
For now, this is the architecture and the use cases d-Matrix is proposing. The announcement doesn’t include independent Raptor benchmarks or figures that would allow comparing its performance, power draw, or cost per token with Vera Rubin or other inference accelerators.
So claims about lower latency or better energy efficiency should be treated as the manufacturer’s goals and statements until final hardware and comparable measurements are available.
NVLink Fusion Turns NVIDIA’s Infrastructure Into a Platform for Other Chips
The strategic part of the deal goes beyond d-Matrix.
For years, NVLink has been one of NVIDIA’s advantages for building systems where many GPUs function as one large compute domain. With NVLink Fusion, the company is opening that interconnect infrastructure to certain third-party processors.
The result is a seemingly contradictory relationship: other chipmakers can compete with NVIDIA’s GPUs on certain workloads while using NVIDIA technologies to build their systems.
For NVIDIA, this widens the market for NVLink, Spectrum-X, BlueField, ConnectX, MGX, and other components even when the main accelerator isn’t its own.
For chip designers, it removes a different kind of barrier.
Building a custom accelerator is already extremely expensive, but turning it into a platform of hundreds or thousands of processors ready for a data center means solving power delivery, liquid cooling, connectivity, servers, racks, software, and manufacturing.
NVLink Fusion lets specialized companies try to enter that market without rebuilding the entire infrastructure NVIDIA has built around its GPUs.
d-Matrix is a good example because its focus sits precisely in inference. Rather than trying to build a full platform that competes component by component with NVIDIA, it wants to place Raptor inside an architecture compatible with that same kind of infrastructure.
If the timeline holds, the first resulting hardware will show up in late 2027. Until then, two big unknowns remain: how Raptor will actually perform once it’s manufactured, and how much of the advantage promised by its memory architecture holds up once it’s deployed at rack scale.
What the deal does show is NVIDIA evolving beyond just selling its own GPUs. With NVLink Fusion, it’s trying to turn its interconnect, networking, and rack designs into shared infrastructure that some of the specialized chips competing for a slice of the growing AI inference market can also run on.

