Everpure, formerly Pure Storage, has announced new capabilities to help enterprises move AI from testing to production: agent access to its data catalog via MCP, an inference accelerator for FlashBlade that the company says cuts time-to-first-token by up to 20x, and continuous data compression. The announcement was made on September 30, 2026 in London, and the new features will roll out during October. The announcement does not include pricing.
Key facts about Everpure’s production AI announcements in 30 seconds
- Everpure Data Intelligence integrates with MCP so agents can query the data catalog in natural language.
- PureKVA preloads context into GPU memory directly from FlashBlade.
- The company claims up to 20x less time-to-first-token, without detailing the tests behind it.
- DeepReduce continuously compresses data on FlashBlade.
- A reference architecture using open models aims to cut spending on third-party model tokens.
The announcements build on the “Data Primacy” concept Everpure introduced in June at its Pure//Accelerate conference: the idea that data, not applications, should be the center of enterprise architecture in the AI era. The company identifies three obstacles to moving AI into production: fragmented context, unpredictable inference costs, and slow, complex deployments. It’s the same thread we covered when explaining how Everpure wants AI to stop stumbling over out-of-context business data.
Data with context for agents
The first part of the announcement centers on Everpure Data Intelligence, the layer that discovers, classifies, and adds context to company information at its source, whether on Everpure’s own platform or across public clouds, SaaS applications, and third-party storage. It’s an area the company strengthened with its acquisition of 1touch, announced when Pure Storage became Everpure.
There are three new features. The first is native integration with the Model Context Protocol (MCP), the open standard that lets AI agents connect to tools and data sources: this way, agents and security tools can query the data catalog in natural language and check each data item’s sensitivity before using it. The second is simplified deployment from the Pure1 console, with no separate management servers or professional-services projects. The third is file analysis that shows who can access each shared resource and how long it’s gone unused, without reading its content, to fix excessive permissions before opening data up to agents.
Prakash Darji, Everpure’s General Manager of Data and Digital Experience, argues that enterprise AI isn’t stalling for lack of good models, but because data isn’t ready for autonomous agents working in real time, and that the goal is for data to be governed, automated, and instantly accessible.
Faster inference, less storage
The second part of the announcement targets performance and cost. PureKVA (Key-Value Accelerator) lets FlashBlade systems preload context directly into GPU memory. In language models, the key-value cache (KV cache) stores calculations already performed on the context, and serving it from storage avoids repeating them. Everpure says this delivers up to 20x less time-to-first-token (TTFT), with multi-tenant support and without moving data. That’s a vendor-stated maximum: the announcement doesn’t specify which models, context sizes, or configurations were measured, and warns that results may vary by environment.
DeepReduce, for its part, is an always-on compression feature that analyzes storage blocks on FlashBlade to find similarities that traditional deduplication misses, even in already-compressed content. According to the company, it expands usable capacity without affecting write performance or requiring manual tuning, though it doesn’t quantify the savings. Rounding out the package is a reference architecture for optimizing token consumption with open-weight models, designed to let enterprises keep control of their data and reduce what they pay third-party model providers.
What’s still unclear
The announcement itself includes a clause common to product announcements: timing and features remain at Everpure’s discretion, and performance metrics are informational, not a commitment. The features are slated for October, but there’s no pricing, no word on whether the new capabilities require additional licenses, and no published tests backing the PureKVA figure or the DeepReduce savings.
The move fits a clear industry trend. Storage vendors are trying to bring AI compute closer to the data and use storage to relieve pressure on GPU memory, one of the most expensive and scarce components in today’s infrastructure. The real value of these features will depend on independent testing and the experience of early production customers.
Frequently asked questions
What is PureKVA?
It’s the Everpure FlashBlade feature that preloads language-model context into GPU memory. According to the company, it cuts time-to-first-token by up to 20x.
What is Everpure’s MCP integration for?
It lets AI agents and security tools query Everpure Data Intelligence’s data catalog in natural language and check each data item’s sensitivity.
When will the new features be available?
Everpure is targeting October 2026, though the announcement notes that timing remains at the company’s discretion.
Is Everpure the same as Pure Storage?
Yes. Pure Storage changed its name to Everpure in 2026 and trades on the New York Stock Exchange under the ticker P.

