Huawei declare Development of the new generation interconnection architecture UnifiedBus (UB). The technology is designed to solve data transfer speed issues in large-scale computing clusters focused on resource-intensive artificial intelligence tasks.
In traditional data center systems, resource efficiency decreases as the platform expands. Huawei said that as a result, in a cluster of 100,000 NPUs, only about 20% of the available computing power is actually used, while the remaining resources are idle waiting for data. Training models with trillions of parameters requires transmitting large amounts of intermediate messages, and modern AI agents generate requests regularly, creating a constant load on the communication lines.
All of this leads to the need to flexibly redistribute changing loads between CPUs and NPUs. In addition, data storage requirements for computing platforms at all levels are increasing dramatically. Therefore, insufficient interconnect bandwidth has become a major bottleneck limiting the performance of modern artificial intelligence clusters.
Image source: Huawei
The UnifiedBus architecture overcomes existing shortcomings and combines more than a dozen different protocols to increase overall interconnect throughput from approximately 100 GB/s to several TB/s while reducing latency (RTT) from 7 μs to 2 μs. The new architecture also provides unified global memory addressing within the AI node Super PoD. In addition, UnifiedBus directly connects CPU, NPU, RAM and SSD, providing decentralized point-to-point communication between them and flexibly allocating available resources.
As part of the UnifiedBus concept, the DDR chip acts as a replacement memory for the NPU, reducing latency for searches, recommendations and other services, and doubling vector search performance. At the same time, as part of the AI accelerator, the requirements for HBM capacity are also reduced. Overall, UnifiedBus, as a global data backbone, has ultra-high throughput and low latency, enabling flexible expansion of computing resources.
Special UnifiedBus LinkBlade interconnect modules feature a cableless design that provides high-speed communications within the rack of a computing cluster. It is said that in a SuperPoD with 4096 NPUs, using LinkBlade can save about 196 kilometers of copper connections. For communication between racks, the UnifiedBus LinkDevice solution provides 176 ports with a throughput of 1.6 Tbps per port, providing 280 Tbps of full-fiber connectivity. In this case, clusters with millions of NPUs can be formed.
Huawei demonstrated the Agent AI super cluster based on UnifiedBus, consisting of SuperPoD TaiShan 950 and Atlas 960 nodes for data storage OceanStor M900 and Galaxy UBG exchanger. The platform supports multi-tier resource pools to enable efficient operation of various workloads.
source:










