China is shifting part of the artificial intelligence battle from individual accelerator performance to the ability to connect thousands of chips to work as a single system. Huawei is the most visible example with Atlas 950 SuperPoD, but they are not alone: Chinese manufacturers like Biren, Moore Threads, Sugon, MetaX, and Enflame are developing similar architectures. The goal is to achieve more effective power even if each individual processor cannot compete directly with Nvidia’s most advanced GPUs.
The key points of Chinese AI supernodes in 20 seconds
- Huawei is preparing Atlas 950 SuperPoD with up to 8,192 NPU Ascend 950DT connected via UnifiedBus.
- China aims to compensate for its limitations in advanced manufacturing by improving interconnection, memory, and scalability.
- Biren, Moore Threads, Sugon, and other manufacturers are also working on supernodes.
- The competition is shifting from isolated chip performance to the performance of the entire system.
- Atlas 950 is scheduled to be commercially available in Q4 2026.
This trend was especially clear during the World Artificial Intelligence Conference (WAIC) held in July in Shanghai. Huawei physically showcased the Atlas 950, while other national manufacturers presented systems based on the same idea: grouping an increasing number of accelerators with high-speed connections and minimizing data transfer costs among them.
It’s a strategy closely related to China’s technological constraints. Chinese companies do not have the same access as their U.S. competitors to the most advanced semiconductor manufacturing processes, production equipment, HBM memory, and packaging technologies.
Huawei has acknowledged this quite explicitly. When unveiling its SuperPoD strategy, Eric Xu, the company’s rotating chairman, explained that faced with limited access to advanced fabrication nodes, one of their solutions is to combine more computing resources.
Atlas 950 aims to turn 8,192 chips into a giant computer
The most prominent example of this philosophy is Huawei’s Atlas 950 SuperPoD.
Its main feature is not just using many processors. Huawei wants up to 8,192 NPU Ascend 950DT to behave as the resources of a single logical machine, rather than simply functioning as thousands of accelerators distributed across a conventional cluster.
That’s where UnifiedBus comes in.
This interconnection technology offers unified memory addressing and high-bandwidth, low-latency communication between different resources. Huawei claims this reduces one of the biggest issues of AI clusters: as the number of accelerators grows, it becomes harder to keep all of them busy with useful computations.
In its full configuration planned for late 2026, Atlas 950 SuperPoD will occupy around 1,000 square meters and consist of 160 racks, with 128 dedicated to computing and 32 to communications.
Huawei announces 8 EFLOPS of FP8 capacity and 16 EFLOPS FP4, along with a combined interconnection bandwidth of approximately 16 PB/s. These figures are provided by the manufacturer and should not be interpreted as direct performance comparisons with other architectures.
The conceptual difference compared to a traditional cluster is even more intriguing than these numbers.
Data centers can host tens of thousands of accelerators yet still face utilization problems if communication among them becomes a bottleneck. Large models require constant exchange of parameters, activations, and other data. Adding chips stops providing proportional improvements once they spend too much time waiting for information.
The supernode aims to extend precisely the domain where accelerators can work in close coordination.
China is no longer only competing by building better GPUs
Huawei is not the only Chinese company reaching this conclusion.
WAIC 2026 showcased how the concept is spreading across the domestic industry. Sugon presented its Dawning 8000, tied to a national supercluster of 100,000 accelerators, while other manufacturers displayed systems with varying numbers of cards.
Biren Technology is betting on optical communications. The company introduced an architecture designed to connect over 1,000 accelerators within a supernode using near-encapsulated optical links to overcome some limitations of traditional electrical connections.
Moore Threads has demonstrated an architecture with 256 GPUs and a single scale-up interconnection layer. Enflame has also collaborated with ZTE on another high-density architecture.
This indicates an interesting shift in the Chinese semiconductor industry.
For years, comparisons with Nvidia focused mainly on processor specs: how many FLOPS an Ascend offers versus a Nvidia GPU, how much memory it has, or what bandwidth it provides.
Supernodes partially change this question.
Now, it also matters how many accelerators can communicate efficiently, what bandwidth exists between them, the latency introduced by the network, how memory is shared, and what percentage of the theoretical capacity is actually converted into tokens or useful training.
This does not mean individual chip performance has become irrelevant. An less efficient accelerator requires more units, more power, more cooling, and potentially more space to accomplish the same task.
But a better-designed system architecture can mitigate some of that disadvantage.
U.S. restrictions are accelerating this architecture
This evolution is also influenced by U.S. controls on exporting advanced technology to China.
Huawei still faces limitations in working with the most advanced manufacturing technologies. Their own supernode strategy is a way to work with manufacturing processes available within China.
The paradox is that restrictions intended to restrict Chinese access to the best chips are forcing manufacturers to devote more resources to networks, packaging, optics, memory, software, and distributed architecture.
This does not eliminate their problems.
TrendForce warns that the Chinese industry continues to face restrictions related to advanced logic manufacturing, high-bandwidth memory (HBM), and advanced packaging. Designing large systems also does not automatically mean having enough chips to fill them.
Huawei plans to produce around 750,000 Ascend 950PRs in 2026, according to Reuters sources, but production might fall short of demand. Companies like ByteDance, Tencent, and Alibaba are interested in acquiring these processors following the launch of DeepSeek V4.
Here lies one of the big contradictions of the Chinese model: designing a supernode for thousands of accelerators is one thing; manufacturing hundreds of thousands of competitive accelerators at scale is another.
DeepSeek helps complete the Nvidia alternative
Hardware alone is not enough without models and software ready to use it.
That’s why the fact that DeepSeek V4 has been adapted to run on Huawei’s Ascend chips is significant. The company also confirmed that its supernodes based on Ascend 950 support different model versions.
It’s arguably a more important piece than it appears.
Nvidia’s dominance isn’t solely due to manufacturing extremely fast GPUs. CUDA, libraries, neural networks, systems like NVL72/NVL144, developer tools, and decades of integration form a comprehensive platform.
China needs to build something similar around its own accelerators.
Huawei is trying this with Ascend hardware, UnifiedBus interconnection, and CANN (Compute Architecture for Neural Networks) as a software layer. The company has open-sourced CANN and is working to improve its integration with projects like PyTorch, Triton, vLLM, and verl.
After SuperPoD, comes the SuperCluster
Atlas 950 is not the ultimate goal for Huawei.
The company aims to connect multiple SuperPoDs to build a Atlas 950 SuperCluster, an infrastructure with over 520,000 Ascend 950DT chips distributed across more than 10,000 racks.
According to their plans, the system will reach 524 EFLOPS FP8 and roughly 1 ZFLOPS FP4. By 2027, they anticipate Atlas 960 SuperPoD, with up to 15,488 NPUs, followed by a SuperCluster with over a million processors.
Again, these are targets and announced specifications, not results from real workloads.
But they illustrate the direction of their architecture.
The Chinese response to restrictions on advanced semiconductors is not only trying to produce a chip like Nvidia’s. It also involves making a much larger number of domestic chips work efficiently as a single machine.
The AI infrastructure battle is shifting from silicon to everything surrounding silicon: memory, optics, interconnections, networking, cooling, software, and the ability to coordinate hundreds of thousands of processors.
And in this domain, China has more room to compete than in a simple GPU comparison.
Frequently Asked Questions
What is Huawei Atlas 950 SuperPoD?
It’s an AI computing architecture designed to connect up to 8,192 NPU Ascend 950DT via UnifiedBus and make them operate as a single logical machine.
When will Atlas 950 SuperPoD be available?
Huawei anticipates that the full configuration of Atlas 950 SuperPoD will be available in the fourth quarter of 2026.
Why is China betting on supernodes?
Because they enable higher combined performance by connecting large numbers of accelerators. This strategy is particularly important given restrictions that make it harder for Chinese companies to access the most advanced manufacturing processes and components.
Can Atlas 950 really compete with Nvidia?
Huawei publishes very favorable comparisons of Atlas 950 against future Nvidia systems, but these are manufacturer estimates. Actual performance will also depend on power consumption, chip utilization, software, reliability, cost, and behavior under real workloads.

