Huawei Unveils OceanStor M900 Context Memory Storage for Scalable KV Cache
Read more
Pandaily
pandaily.com

Huawei Unveils OceanStor M900 Context Memory Storage for Scalable KV Cache

Huawei introduced the OceanStor M900 Context Memory Storage data storage system at the HUAWEI CONNECT 2026 event in Shanghai on September 17. This solution is designed as a storage layer specifically for AI inference SuperPoD workloads, rather than being another accelerator card.

During the presentation, the company focused on product specifications related to the KV cache infrastructure for artificial intelligence. The roadmaps for Ascend SuperPoD accelerators and the Peerium architecture components presented during the same period were not discussed.

As models become trillion-parameter and context windows exceed one million tokens, multi-turn agent workloads cause the KV cache to exceed onboard memory and DRAM. Huawei's answer is a globally unified, multi-level KV cache supported by UnifiedBus, which extends SuperPoD memory from SRAM and DRAM to SSD.

According to company materials, one cluster can achieve a capacity of 64 petabytes. This allows increasing the available KV cache per NPU from gigabytes to terabytes for reuse, thereby improving cache hit rates during long-context inference.

The system's performance is based on an integrated architecture combining a Central Processing Unit (CPU), network controller, and NAND controller. This enables SuperPoD NPUs to access SSDs directly, bypassing protocol conversion or transmission through the CPU. Huawei claims that this path reduces access latency from milliseconds to approximately 60 microseconds, representing a claimed reduction of 90%, and provides a cumulative access throughput of about 40 TB/s, which is 1.5 times higher than competing solutions.

In typical AI programming scenarios, the company states that this design can double the token throughput of an inference cluster and halve the time to first token (TTFT); these figures remain vendor claims.

Pricing considerations are addressed through KV-sensitive adaptive storage, which predicts the cache lifecycle value and places data across different media tiers. Huawei points to a rating of up to 24 drive writes per day (DWPD), which is a claimed 16x increase in SSD endurance, as well as three-year stability targets aimed at reducing replacement costs and the cost per token for hyperscale inference. Thus, OceanStor M900 represents a specific Huawei product that integrates compute, networking, and storage for a petabyte-scale general context cache, rather than just a chip launch.

Similar stories

Huawei introduces Lingqu UnifiedBus as the foundation for the Agentic SuperPoD cluster architecture
Read more
pandaily.com

Huawei introduces Lingqu UnifiedBus as the foundation for the Agentic SuperPoD cluster architecture

At the Huawei Connect 2026 event in Shanghai on September 17, Huawei ICT BG CEO Yan Chaobin introduced Lingqu UnifiedBus as the central interconnection element for a new joint cluster and SuperPoD architecture designed for Agentic AI workloads.

Huawei explained that as clusters grow, efficiency often decreases due to card communication latency. Furthermore, training models with 10 trillion parameters and frequent data exchange between agents and models push the KV Cache and intermediate data far beyond the memory of a single card, making the factory design a performance bottleneck.

The core principles of Lingqu's design include protocol unification: over ten interconnect protocols are merged into a single factory with Lingqu memory semantics. Interconnect bandwidth increases from the 100 GB class to the TB class, and round-trip time is reduced from approximately 7 microseconds to 2 microseconds, enabling global memory access within the SuperPoD.

Central Processing Units (CPUs), Neural Processing Units (NPUs), memory, and Solid State Drives (SSDs) are connected as equal elements for decentralized access. A flexible ratio of CPU to NPU is provided, along with hardware acceleration for Attention and FFN splitting for AF-distributed deployment. Multi-level storage pools can handle activations and use DDR as secondary memory for NPUs, while the optical network is positioned as a high-speed, low-latency data transmission channel for elastic scaling and expansion.

Huawei also announced multi-level Lingqu interconnect hardware, covering rack, inter-rack, and cluster levels. Modules within the rack eliminate losses from copper cables and circuitry—Huawei claims that a 4096-card SuperPoD can save about 196 kilometers of copper cable. Inter-rack switches provide 176 ports with 1.6 Tbps bandwidth each, delivering 280 TB of optical bandwidth per chassis at an RTT delay of about 2 microseconds. The Lingqu Xinghe UBG network switches advertise a branching factor of 1024, designed to support the construction of SuperClusters with millions of cards as models grow to tens of trillions of parameters.

At this factory, Huawei described the Agentic SuperPoD cluster, which combines Kunpeng 950, Ascend 960 SuperPoD, OceanStor M900 memory storage, and UBG switches for heterogeneous computing and unified resources. A similar interconnection story extends to devices: the Atlas 650E air-cooled server can directly connect 16 NPUs across two nodes without a switch, functioning as a small SuperPoD for on-site trillion-parameter inference. The emphasis is placed on co-designing the system and interconnects, not just the near-packet optics of the Ascend 960 SuperPoD, and this remains a vendor roadmap statement until independent cluster measurements are available.

Huawei Unveils Ascend 960 SuperPoD with Near-Packaged Optics at Connect 2026 Conference
Read more
pandaily.com

Huawei Unveils Ascend 960 SuperPoD with Near-Packaged Optics at Connect 2026 Conference

At the Huawei Connect 2026 event, held in Shanghai on September 17th, Chairman David Wang introduced the Ascend 960 SuperPoD as the next-generation supernode for artificial intelligence. The development focus is placed on interconnect scalability rather than single-chip performance.

According to the system description, the Ascend 960 SuperPoD is the first model to utilize Near-Packaged Optics (NPO). This technology integrates Huawei's Lingqu UnifiedBus factory with the Hi-ONE optical engine, enabling optical connections to be placed closer to the chip package.

According to First Finance and related reports, one Ascend 960 SuperPoD can connect approximately 4096 cards with a round-trip latency approaching 2 microseconds. Huawei asserts that these metrics allow for large-scale training and inference of models up to 10 trillion parameters.

Furthermore, broader delivery information was disclosed: Ascend SuperPoD systems have already been shipped to over 1000 customers from more than 370 companies, indicating the architecture's transition from demonstration stands to commercial use.

The chip roadmap has been adjusted. Huawei announced that the Ascend 960DT is planned for release in the first quarter of 2027, and the Ascend 960PR in the third quarter. The liquid Atlas 960 SuperPoD is also scheduled for release in the third quarter of 2027, three quarters earlier than the initial public forecast, which pointed to the end of 2027 for the Ascend 960. Wang also stated that Huawei has developed 11 UnifiedBus-based chips for large systems and confirmed the long-term goal of creating a million-card SuperCluster.

The main emphasis is on system engineering: this includes NPO optics, UnifiedBus, cooling systems, and cluster software, which allow multiple Ascend cards to function as a single machine. Reuters reports confirm the dual 2027 launch dates and the figures of over 1000 supernodes and more than 370 customers, but did not name the buyers. This news differs from Huawei's 'Intelligent World 2035' report published recently, as it presents specific information about the SuperPoD interconnect and the revised Atlas 960 schedule for operators who are already evaluating Ascend clusters.

Popular