Huawei introduced the OceanStor M900 Context Memory Storage data storage system at the HUAWEI CONNECT 2026 event in Shanghai on September 17. This solution is designed as a storage layer specifically for AI inference SuperPoD workloads, rather than being another accelerator card.
During the presentation, the company focused on product specifications related to the KV cache infrastructure for artificial intelligence. The roadmaps for Ascend SuperPoD accelerators and the Peerium architecture components presented during the same period were not discussed.
As models become trillion-parameter and context windows exceed one million tokens, multi-turn agent workloads cause the KV cache to exceed onboard memory and DRAM. Huawei's answer is a globally unified, multi-level KV cache supported by UnifiedBus, which extends SuperPoD memory from SRAM and DRAM to SSD.
According to company materials, one cluster can achieve a capacity of 64 petabytes. This allows increasing the available KV cache per NPU from gigabytes to terabytes for reuse, thereby improving cache hit rates during long-context inference.
The system's performance is based on an integrated architecture combining a Central Processing Unit (CPU), network controller, and NAND controller. This enables SuperPoD NPUs to access SSDs directly, bypassing protocol conversion or transmission through the CPU. Huawei claims that this path reduces access latency from milliseconds to approximately 60 microseconds, representing a claimed reduction of 90%, and provides a cumulative access throughput of about 40 TB/s, which is 1.5 times higher than competing solutions.
In typical AI programming scenarios, the company states that this design can double the token throughput of an inference cluster and halve the time to first token (TTFT); these figures remain vendor claims.
Pricing considerations are addressed through KV-sensitive adaptive storage, which predicts the cache lifecycle value and places data across different media tiers. Huawei points to a rating of up to 24 drive writes per day (DWPD), which is a claimed 16x increase in SSD endurance, as well as three-year stability targets aimed at reducing replacement costs and the cost per token for hyperscale inference. Thus, OceanStor M900 represents a specific Huawei product that integrates compute, networking, and storage for a petabyte-scale general context cache, rather than just a chip launch.


