NetApp Activates Enterprise Data Without Copies to Get Ready for Agentic AI

NetApp has unveiled new capabilities that let companies use their data directly where it already sits, without having to create dedicated copies to feed artificial intelligence applications. The update centers on NetApp AI Data Engine (AIDE), which extends metadata discovery across heterogeneous environments, and on a strengthened integration with Commvault to detect threats and speed up the recovery of data used in AI workloads.

Agentic AI and enterprise data: the key points in 20 seconds

  • NetApp extends AI Data Engine to discover metadata across both NetApp and third-party storage.
  • The goal is to activate existing data without copying it to new platforms for every AI project.
  • AIDE works with NFS, SMB, and S3 repositories and adds governance and data-understanding capabilities.
  • The Commvault integration connects ransomware detection with clean recovery.
  • Some of the announced features are still in preview and may change before they become available.

The pitch starts from a problem that’s becoming more visible as AI agents move from testing into production environments. Companies accumulate information across local storage, cloud services, objects, files, and systems with different access and protection policies. Preparing that data for new models usually means moving it, transforming it, or creating additional copies.

NetApp wants to cut down on exactly that dependence on copies. Its approach adds a layer of discovery, understanding, and governance on top of existing data so it can be used in analytics applications, assistants, and AI agents without necessarily moving it to another repository.

The company presents this approach as a zero-copy model — that is, activating data where it already sits instead of creating a new copy as a preliminary step for every AI application. It’s a similar logic to the one behind Teradata’s recent move to query AI workloads directly on data stored in Microsoft OneLake without copying it first, applied here to NetApp’s own storage layer.

AI Data Engine extends discovery to data across different systems

AI Data Engine is part of NetApp’s data services layer for AI. With this update, the company extends metadata discovery to cover its own storage as well as repositories that don’t run NetApp technology.

The tool can discover and analyze information stored via NFS, SMB, and S3. NFS (Network File System) and SMB (Server Message Block) are common protocols for file access, while S3 is mainly used for object storage.

The stated goal is for organizations to be able to locate information that until now has stayed scattered across different systems and determine which data can be useful for AI projects. NetApp also ties these capabilities to tasks such as identifying security risks, optimizing storage, easing migrations, and finding information that’s already AI-ready.

The update also adds new capabilities that are still in preview. These include AI-driven data-understanding features, expanded governance, integrations with analytics engines such as Starburst and Onehouse.AI, and new agent-oriented services.

NetApp also notes that this approach could serve as a foundation for later capabilities such as semantic search, knowledge graphs, and context generation for agents. Since these are future or preview features, the company warns that their availability, characteristics, and timing may change.

The difference from a traditional architecture lies in where the work happens. Rather than first building a new repository with a copy of enterprise information, the approach aims to have the storage infrastructure itself provide context about the data it already manages.

For companies, this can reduce some of the processes tied to moving information around, although the actual benefit will depend on the systems in use, permissions, and the requirements of each environment.

Commvault adds attack recovery to the AI data strategy

The second part of the announcement focuses on protecting that data. NetApp has expanded its integration with Commvault to connect its platform’s ransomware-protection signals with Commvault’s recovery workflows, in the same spirit as Commvault’s own standby Active Directory service for recovering identities in minutes.

The solution combines threat detection at the primary storage layer with processes designed to find and recover clean data. According to NetApp, it includes near-real-time threat detection, deeper analysis to spot attacks that try to evade protection systems, and data validation in an isolated environment before recovery proceeds.

The integration is managed from the Commvault console and aims to cut the time needed to recover information after an attack. NetApp says the goal is to move from recovery processes that can take days or weeks to recoveries measured in minutes in certain scenarios.

That last claim describes the solution’s pitch and shouldn’t be read as a guaranteed recovery time for any infrastructure. The actual result will depend on the environment, the volume of data, the type of incident, and how the systems are configured.

The link between protection and recovery becomes more important as the same data starts feeding AI applications. A compromised or altered data set can affect not only traditional applications but also models, agents, and automated processes that use that data as context.

That’s why NetApp frames both announcements as part of the same strategy: preparing data to be used by AI systems without separating that preparation from existing governance and protection policies.

For agents, the challenge isn’t just accessing the data

The arrival of AI agents introduces a difference compared with traditional assistants. An agent can query information, combine data from different sources, and execute actions at a speed far higher than a manual workflow.

That makes access to data insufficient on its own. It’s also necessary to know which information each agent can use, what permissions it has, whether the data is current, and what mechanisms exist to recover a reliable version when an incident occurs.

NetApp positions AI Data Engine precisely around that combination of discovery, understanding, governance, and activation. The company argues that enterprise data can become context for agents without necessarily leaving the systems where it’s already stored.

Shari Lava, vice president and general manager of AI research at IDC, says agents need current, reliable, and governed data, and believes that connecting discovery, understanding, governance, activation, and resilience from the storage layer can give companies a way to prepare their data without facing major data-movement projects.

That statement comes from IDC and is part of the assessment NetApp included in its announcement, so it should be read as the analyst’s opinion rather than an independent validation of all the commercial capabilities presented.

The company also insists that its approach isn’t meant to become a closed AI platform. The announcement explicitly mentions open integrations with analytics engines and the ability to work across heterogeneous infrastructures.

That point matters because today’s enterprise architectures are rarely built around a single storage system. Data can be spread across data centers, public clouds, file systems, object storage, and sovereign environments.

NetApp’s goal is for that diversity to not force companies to build a different data copy every time a new AI application shows up.

What remains an open question is how broadly this model can be applied. Some of the announced capabilities are available as new features, while others are still described as planned developments or preview releases. NetApp itself warns that future features, their timelines, and availability conditions may change.

For companies moving from experimenting with AI to using agents in real processes, the move points to an issue less visible than the models themselves: the quality, location, governance, and recoverability of the data agents need to act.

Frequently asked questions

What is NetApp AI Data Engine?

AI Data Engine is a NetApp service aimed at discovering, understanding, governing, and preparing enterprise data for analytics, assistants, and AI applications. The new update extends metadata discovery to both NetApp and third-party storage.

What does zero-copy mean in NetApp’s approach?

The zero-copy approach aims to activate data where it’s already stored without creating another dedicated copy for each AI project. The company positions this model as an alternative to architectures that move and transform data first.

What does the NetApp-Commvault integration add?

The integration connects NetApp’s ransomware-protection signals with Commvault’s recovery processes. It includes threat detection, analysis, validation in isolated environments, and recovery of clean data.

Are all of AI Data Engine’s new features already available?

No. NetApp is announcing some capabilities as new features while presenting others as previews or future developments. The company warns that the features, availability, pricing, and timelines of unreleased offerings may change.

Scroll to Top