Huawei has introduced OceanStor M900 Context Memory Storage, a storage system designed to accelerate AI inference in large-scale data centers. The proposal extends the memory available for KV cache, used during large-model inference, from accelerator memory and DRAM to SSD units, with an aggregate capacity Huawei puts at up to 64 PB per cluster.
OceanStor M900 at a glance
- Huawei extends SuperPoD KV cache up to 64 PB per cluster.
- The system combines accelerator memory, DRAM, and SSD through UnifiedBus.
- Huawei says it cuts access latency from milliseconds to 60 microseconds.
- The vendor claims up to 40 TB/s of aggregate bandwidth and double the token throughput.
- An adaptive storage technology aims to extend SSD lifespan by 16 times.
Huawei unveiled OceanStor M900 on September 17 during HUAWEI CONNECT 2026, held in Shanghai. The system is part of a strategy in which the company is trying to extend SuperPoD architecture beyond compute and tightly coordinate compute, networking, and storage.
The reasoning lies in how AI models are evolving. As context windows grow and agents keep conversations and tasks running longer, the amount of data that needs to be retained during inference also increases. Huawei identifies KV cache as one of the elements starting to strain the capacity and cost of the memory available on accelerators and in DRAM.
The company positions OceanStor M900 as an additional layer for storing and reusing that information without relying solely on the memory sitting next to the processing units.
A KV cache that reaches all the way to SSDs
The KV cache stores the Key and Value data generated during a model’s inference. When part of the context can be reused, the system can retrieve that data instead of running certain calculations again.
This mechanism becomes more important with models capable of working with context windows beyond a million tokens, according to Huawei. Multi-turn conversations and complex tasks also generate more information that can remain available for later operations.
The problem is that accelerator memory and DRAM have limited capacity and a cost that rises when scaling to large volumes is needed.
OceanStor M900 uses UnifiedBus, Huawei’s interconnect network, to create a global, tiered KV cache space. Information can be distributed across accelerator memory, DRAM, and SSD, so total available capacity grows without depending exclusively on the memory closest to the processor.
Huawei says a single cluster can reach 64 PB of KV cache capacity. It also notes that capacity available per NPU goes from gigabytes to terabytes.
The company argues that this capacity increase makes it possible to store, share, and reuse more context and can raise the cache hit rate. These figures and effects are claims made by the manufacturer, and the published information doesn’t include independent validation.
From milliseconds to 60 microseconds, according to Huawei
The second part of the proposal relates to how stored data is accessed.
Huawei says OceanStor M900 incorporates an architecture that integrates a CPU, a network controller unit, and a NAND controller unit. The design provides native semantics for the KV cache and enables a single-hop connection between the SuperPoD’s NPUs and the SSDs.
According to Huawei, this architecture avoids protocol conversions and certain CPU-mediated transfer operations. The manufacturer says it cuts access latency from the millisecond range down to 60 microseconds, which by its calculation amounts to a 90% reduction.
The system also reaches, according to Huawei, up to 40 TB/s of aggregate bandwidth per cluster, a figure the company puts at 1.5 times above comparable solutions.
The intent is to reduce the time accelerators spend waiting for data. In AI workloads where computation and context access are closely tied, more storage capacity doesn’t help much if retrieving that information introduces significant wait time.
Huawei says that, in typical AI scheduling scenarios, OceanStor M900 can double the token throughput of the inference cluster and cut time to first token (TTFT) in half. These are, again, results reported by Huawei in its own benchmark scenarios.
Storage also has to withstand wear
The large volume of writes a caching infrastructure can generate raises another problem: SSD endurance.
To address this, Huawei includes what it calls KV-aware adaptive storage technology, a technology that factors in the behavior and lifecycle of KV cache data to decide how to distribute it across the different storage media.
The idea is to identify the value and expected duration of the data and adapt where it’s placed. That way, not all writes concentrate the same way on the same devices.
Huawei states that OceanStor M900 can reach up to 24 drive writes per day (DWPD) and that this technology can multiply SSD endurance by 16 times. The company puts the expected stability of the drives at three years under the described conditions.
The goal is to reduce both media replacement and the operating and maintenance costs tied to large inference infrastructures.
From a compute-centric infrastructure to one built on three layers
The M900 launch fits into a broader Huawei strategy for its AI systems. The company believes that the growth of models and agents is giving context storage more weight within the infrastructure.
Until now, a significant part of AI system design has focused on increasing GPU or NPU compute capacity. Huawei is now proposing an architecture in which compute, networking, and storage need to work in a coordinated way, an approach that follows the roadmap the company laid out with its Ascend 960 chips and Atlas 960E SuperPoD.
In large SuperPoDs, the volume of information moving between accelerators and memory can become a limitation. A distributed, shared KV cache, according to Huawei’s proposal, makes it possible to retain more context and reuse it from different compute resources.
This is especially relevant for inference systems that maintain long contexts or handle multi-step tasks. In these scenarios, part of the work consists precisely of retrieving previously generated information and making it available to the model at the right moment.
OceanStor M900 aims to turn SSDs into an extension of that context memory, though with a different capacity and access hierarchy compared with memory located on the accelerator itself.
The proposal is part of the infrastructure announcements Huawei made at HUAWEI CONNECT 2026 alongside its new SuperPoD and SuperCluster architectures, following the same direction it set with the CloudMatrix 384 AI cluster. Overall, the company is proposing an infrastructure in which AI performance doesn’t depend solely on how many processing units a system can gather, but also on how quickly those units can access the context they need.
Frequently Asked Questions
What is OceanStor M900?
It’s a Huawei storage system designed for AI inference in large-scale data centers. Its main function is to extend and share KV cache across accelerator memory, DRAM, and SSD.
What capacity can it reach?
Huawei says a cluster with OceanStor M900 can provide up to 64 PB of KV cache capacity.
What is KV cache?
It’s a structure that stores data generated during a model’s inference so it can be reused later, avoiding certain repeated calculations. Its size grows with long context windows and multi-turn tasks.
What performance improvements does Huawei claim?
The company says the system can cut access latency to 60 microseconds, reach 40 TB/s of aggregate bandwidth, double token throughput, and cut time to first token in half in certain scenarios.

