Nvidia and Broadcom Clash Over the Future of AI Infrastructure

The growth of artificial intelligence is opening a new battleground in data centers. It’s no longer just about who makes the most powerful accelerator, but about how to connect tens of thousands of chips and scale computing beyond a single rack. Nvidia and Broadcom represent two different strategies to solve this challenge: programmable, platform-based solutions versus custom chips and networks for hyperscalers.

Nvidia aims to provide virtually the entire infrastructure needed to build an AI Factory, from GPUs and CPUs to interconnects, switches, and software. Broadcom, on the other hand, is expanding through custom accelerators—often called XPUs—and networking technologies developed for a select group of giant tech companies handling volumes that justify their own designs.

Key points in 30 seconds

  • Nvidia champions a programmable platform available to multiple types of clients.
  • Broadcom develops custom accelerators alongside large hyperscalers.
  • ASICs/XPU can optimize performance, energy efficiency, and cost for very specific workloads.
  • GPUs offer greater flexibility when models and algorithms evolve.
  • Networking is becoming a fundamental part of performance for large AI clusters.
  • Nvidia controls much of its stack via CUDA, GPU, NVLink, InfiniBand, and Ethernet.
  • Broadcom has a particularly strong position in Ethernet switching and custom silicon.
  • The next battleground is connecting thousands of accelerators as if they were a massive computing system.
  • Hyperscalers likely will combine general-purpose GPUs with their own accelerators instead of opting solely for one architecture.

Two different ways to build AI infrastructure

The difference between Nvidia and Broadcom starts with their very business models.

Nvidia has built a horizontal ecosystem around its GPUs. Its accelerators can end up in AWS, Microsoft Azure, Google Cloud, Oracle Cloud, specialized providers, supercomputing centers, sovereign clouds, or enterprise private infrastructure.

But selling GPUs is only part of the strategy.

Nvidia has been evolving into a full infrastructure company for years. CUDA provides the software layer, while technologies like NVLink, NVSwitch, InfiniBand, and Ethernet platforms enable the connection of increasing numbers of accelerators.

Broadcom operates from a different position.

The company has become a major player in ASICs and custom accelerators for AI, working directly with some of the world’s largest infrastructure operators.

Instead of selling exactly the same accelerator to thousands of companies, it collaborates with hyperscalers to develop silicon specifically optimized for their needs.

Programmable GPUs versus custom XPUs

This difference is fundamental to understanding the battle.

A Nvidia GPU is a programmable platform.

It can be used to train large language models, perform inference, work with computer vision, develop scientific models, or test new architectures that did not exist when the chip was fabricated.

An ASIC, on the other hand, is designed with a much more specific goal.

When a company has a perfect understanding of the operations it will run at massive scale, it can design hardware that eliminates unnecessary components and optimizes others.

The potential result is a significant improvement in cost, energy consumption, and performance for that particular workload.

That’s why custom chips make particular sense for companies like Google, Meta, or large AI model developers.

When deploying hundreds of thousands of accelerators, even a small efficiency gain can lead to enormous savings.

Nvidia’s advantage is flexibility

Nvidia’s main defense against custom silicon is precisely the speed at which AI evolves.

Models are evolving extraordinarily fast.

New architectures, attention techniques, numeric formats, mixture-of-experts systems, and inference methods can change computational needs in a few months.

Designing an ASIC takes a considerably longer cycle: defining architecture, chip design, validation, fabrication, and deployment.

A programmable platform can adapt through software to many of these changes.

This flexibility accounts for much of the GPU success in AI.

A company doesn’t need to know exactly which models it will run in three years to buy Nvidia infrastructure today.

Broadcom’s other advantage: hyperscalers know their needs

But the argument works in reverse too.

The biggest cloud operators run workloads at a scale that few other organizations can match.

When a hyperscaler knows that a particular type of inference will be executed billions of times, optimizing hardware specifically for that makes economic sense.

Google has been demonstrating this approach with its TPU chips for years.

Meta develops its own MTIA accelerators. AWS offers Trainium and Inferentia. Microsoft works on its own silicon with Maia.

The growth of custom silicon shows that major GPU buyers also want to avoid reliance on a single architecture indefinitely.

And here, Broadcom sees a huge opportunity.

The battle is no longer just inside servers

There’s also another challenge that’s completely changing the market.

A standalone accelerator has limited capacity.

The largest models need to work across hundreds, thousands, or even tens of thousands of accelerators simultaneously.

Therefore, system performance increasingly depends on how these chips communicate.

The network ceases to be just the mechanism to connect servers and becomes part of the computing system itself.

Here, two fundamental concepts emerge: scale-up and scale-out.

Scale-up: making many GPUs work as a single system

Scale-up involves, simplistically, connecting accelerators via extremely fast and low-latency links within a tightly integrated compute domain.

Nvidia has established a strong position with NVLink and NVSwitch.

Their rack architectures enable connecting multiple GPUs with enormous bandwidth, making the whole behave more like a single large accelerator.

This approach reaches a new level with rack-scale systems, where an entire rack begins to serve as the basic computing unit.

But the largest clusters need to go even further.

Scale-out: connecting thousands of accelerators

When many racks are connected, scale-out networking comes into play.

Technologies like InfiniBand and Ethernet enable fabrics that connect thousands of nodes.

Nvidia has a significant position here thanks to its acquisition of Mellanox, which provided key high-performance networking technologies.

However, Ethernet is advancing rapidly, and Broadcom is one of its leading silicon suppliers.

As hyperscalers aim to build ever-larger clusters with open architectures, Ethernet could gain even more prominence.

Broadcom wants Ethernet to scale AI

Broadcom’s importance isn’t only in its custom accelerators.

The company is a major player in data center switching.

Chip families like Tomahawk are part of many cloud infrastructures, enabling the construction of Ethernet networks with increasing capacity.

Strategically, the question is whether Ethernet can increasingly handle the communication needs of gigantic AI Factories.

If so, Broadcom could benefit regardless of who makes the accelerators.

A cluster can use Nvidia GPUs while simultaneously incorporating Broadcom-based networking technologies.

That’s why the relationship between these companies isn’t simply a direct competition.

Nvidia aims to control the entire stack

Nvidia’s strategy is precisely to close those gaps where others could enter.

It no longer wants to just sell the chip.

Its offering includes:

GPU + CPU + interconnect + networking + systems + software.

CUDA is probably the hardest element to replicate.

Thousands of applications, frameworks, and tools are optimized for Nvidia’s ecosystem.

There’s also increasing integration of hardware and networking.

This allows Nvidia to optimize the entire system rather than individual components.

For customers wanting to deploy AI infrastructure quickly, buying an end-to-end validated architecture can be much easier than integrating technologies from multiple providers.

Broadcom bets on customization

Broadcom plays a different game.

Its large clients have thousands of engineers and can afford to design their own architecture.

They don’t necessarily need to buy a complete system.

They can choose their preferred accelerator, network architecture, and how everything integrates into their data centers.

Controlling the design might be more important for them than buying a ready-made platform.

Economic incentives are significant too.

At multi-gigawatt scales, reducing energy consumption or silicon costs by just a few percentage points can save hundreds of millions of dollars.

Energy consumption accelerating change

Power availability is becoming one of the biggest constraints on AI growth.

New campuses are projected at hundreds of megawatts, and some developments are moving toward gigawatt scales.

In this context, efficiency per watt is as important as total performance.

Custom accelerators can be particularly attractive for inference workloads where certain tasks are more predictable and repetitive.

GPUs retain a clear advantage when flexibility and running varied workloads are priorities.

This could lead to a natural division.

The future will likely be hybrid

The Nvidia-Broadcom battle doesn’t necessarily end with a single winner.

Large AI data centers will probably be heterogeneous.

Programmable GPUs may continue to dominate certain training, research, and rapidly evolving workloads.

Custom accelerators can gradually take on large volumes of inference and well-understood workloads.

CPUs will manage other parts of the system.

Different network technologies will connect all this infrastructure.

The key question is what percentage of spending will go to each layer.

The biggest risk for Nvidia isn’t another GPU

For years, the competitor to Nvidia was looked for among manufacturers capable of developing a better GPU.

But the most significant threat may come from a different direction.

It’s not necessary to replace Nvidia GPUs entirely.

It’s enough for major clients to shift an increasing share of their workloads to their own accelerators.

If Google, Meta, Amazon, Microsoft, and others progressively use more custom silicon, the market will continue to grow even as the percentage passing through a commercial GPU declines.

Broadcom can benefit directly from that transition by providing design, IP, interconnect, and networking solutions.

But Nvidia has something extraordinarily hard to replicate

Nvidia also doesn’t compete solely on hardware specs.

It competes with an ecosystem built over decades.

CUDA, optimized libraries, development tools, frameworks, networking, and complete systems greatly reduce the difficulty of deploying new AI capacity.

This advantage is especially important outside the small group of hyperscalers capable of designing their own chips.

A company, government, university, or regional cloud provider is unlikely to develop their own AI ASIC.

For these clients, a programmable platform still makes much more sense.

The next AI war will be over data center architecture

Nvidia and Broadcom thus represent two different visions of the future.

Nvidia seeks to make the data center a giant platform-defined computing machine, where accelerators, networks, and software evolve together.

Broadcom advocates for an infrastructure where the biggest operators can build customized systems combining their own accelerators with increasingly fast Ethernet networks.

Both strategies can coexist.

But one thing is becoming clear: the next phase of the AI race will be decided less by FLOPS alone and more by overall system architecture.

When clusters contain tens of thousands of accelerators, networking, interconnects, energy efficiency, and holistic data center design are as critical as the processor.

The battle isn’t just about manufacturing the fastest GPU anymore.

It’s about deciding how to build the massive machines that will run the next generation of AI.

Scroll to Top