Huawei Atlas 960E scales up to a SuperCluster of a million NPUs

Huawei Atlas 960E scales up to a SuperCluster of a million NPUs

Huawei is accelerating its Ascend roadmap and introducing the Atlas 960E SuperPoD. The company touts it as the first system based on near-packaged optics. The SuperCluster scales up to one million NPUs.

David Wang, Deputy Chairman of the Board and Rotating Chairman, outlined the strategy at the Huawei Connect 2026 conference in Shanghai. The core of his message: AI infrastructure should not be a collection of disparate servers, but a highly integrated system. Huawei is building this vision around SuperPoDs and SuperClusters, using its proprietary UnifiedBus interconnect as the binding element.

The figures Wang cites to support his argument are staggering. Foundation models are approaching ten trillion parameters and are projected to exceed one hundred trillion by 2030. In China alone, some 500 trillion inference tokens are processed daily. And on-device models for smartphones have grown from three billion parameters in 2024 to thirty billion today.

One generation per year

The SuperPoDs run on Ascend, the NPUs that Huawei builds itself. “We’re evolving our Ascend chip series on a one-generation-a-year cycle,” said Wang. The Atlas 960E SuperPoD will eventually run on the Ascend 960DT, which is set to be released in the first quarter of 2027. The Atlas 960E SuperPoD will take a little longer to arrive than the server chip itself, specifically, until the third quarter of next year.

The Ascend 960DT offers up to 2 PFLOPs of FP8 computing power and 4 PFLOPs of FP4 computing power. This makes the NPUs ideal for demanding AI workloads. The chips also feature up to 288 GB of High Bandwidth Memory, operating at a speed of 9.6 TB/s. That’s 2.4 times faster than the previous Ascend generation. The inter-chip bandwidth is 2.2 TB/s per chip.

The Ascend 970 and 980 are expected to follow in 2028 and 2029, respectively. Thanks to what Huawei calls the Tau scaling law, computing specifications, memory bandwidth, memory capacity, and interconnect bandwidth are expected to grow significantly.

Why SuperPoDs?

In the SuperPoD system, multiple nodes are tightly integrated via high-speed interconnect protocols, with unified memory addressing across physical nodes. The entire system behaves as a single logical computer.

Clusters of 100,000 NPUs have become the standard for training state-of-the-art models. In traditional server architectures, communication within the cluster consumes well over 40 percent of the total training time, which significantly limits Model FLOPs Utilization. Simulations by Huawei’s Markov Lab show that a cluster of 100,000 NPUs built from SuperPoDs of 4,000 NPUs achieves a 2.75 times higher MFU than the same cluster based on servers with eight NPUs.

Optics closer to the chip

The most interesting hardware announcement in optics. Huawei is introducing the High-density Optical Interconnect Node Engine, or Hi-ONE for short, based on near-packaged optics (NPO). A single engine achieves a transmission capacity of 7.2 Tbit/s. Huawei states that the product is the first NPO component ready for mass production and the only one with a built-in light source.

NPO places the optical engine in close proximity to the application-specific integrated circuit (ASIC). This results in lower power consumption, reduced signal loss, and improved cooling. It also offers modularity and maintainability.

Using Ascend 960 chips and Hi-ONE, Huawei is building the Atlas 960E SuperPoD. It scales to 4,096 NPUs, delivers 8 EFLOPS of FP8 computing power, and offers up to one petabyte of HBM capacity. By deploying 5,500 Hi-ONE units, the 48,000 800G optical modules normally required are no longer needed. This saves more than 550 kilowatts of power. Uptime doubles, with system availability reaching 99.8 percent.

TaiShan, OceanStor, and a million NPUs

Beyond AI, Huawei is also updating the general-purpose TaiShan 950 SuperPoD. It now supports up to 4,096 nodes with a unified memory pool of up to 256 TB. For sandbox-intensive workloads, 100,000 sandboxes launch 30 times faster than on traditional servers, while density increases by a quarter. Vector queries across ten billion thousand-dimensional vectors are twice as efficient.

In addition, OceanStor M900 is a context memory storage cluster on UnifiedBus. The system provides multi-tier KV caching for agentic inference, with one-hop direct access and a petabyte-scale KV cache for the L3.5 layer. Through hybrid media and an optimized retention algorithm, Huawei extends the read and write lifespan of SSDs sixteenfold.

Taken together, this forms the agentic SuperCluster. UnifiedBus consolidates multiple interconnect protocols into a single protocol, reducing protocol-conversion overhead and enabling peer-to-peer connections between Ascend SuperPoDs, Kunpeng SuperPoDs, and KV-cache clusters. With a two-layer, four-plane Clos architecture, the SuperCluster connects up to 512,000 NPUs. Combined with a multi-rail topology, this scales up to one million NPUs.

Ecosystem and developers

Huawei also shared figures regarding the software side. The Kunpeng ecosystem includes more than 4.16 million developers and over 7,200 partners, with support for more than 560 open-source projects. openEuler surpassed 20 million installations, giving it the largest share of the Chinese server operating system market.

For Ascend, the focus is on CANN, the Compute Architecture for Neural Networks. External developers now account for 61 percent of all CANN developers, surpassing the internal group for the first time. The community has over 5,200 monthly active developers. More than forty models have been natively pre-trained on Ascend and CANN. Ascend supports over ninety external open-source projects, including PyTorch, Triton, vLLM, and veRL, and, with support from the Linux Foundation, is the first Chinese compute platform to be officially listed on the PyTorch website.

At Connect, Huawei also announced plans to accelerate the development of Ascend chips for SuperPoD systems. The Ascend 960DT will be available in the first quarter of 2027, three quarters earlier than indicated in the original roadmap. The Ascend 960PR, which offers slightly better performance, will follow in the third quarter of 2027, one quarter earlier.

Tip: Huawei unveils full-stack AI data center strategy