NetApp has acquired DataPelago, a California-based company specializing in accelerating data processing through central processing units (CPUs) and graphics processing units (GPUs). The transaction will enable the manufacturer to incorporate these capabilities into its storage infrastructure and move toward a model where information is prepared for artificial intelligence (AI) and analytics without the need to copy it to another system beforehand.
The key points of the DataPelago acquisition in 30 seconds
- NetApp has acquired DataPelago, though the purchase price and financial terms have not been disclosed.
- Their Nucleus technology processes data using CPUs and GPUs close to where the data is stored.
- The goal is to reduce copying and data movement prior to AI projects.
- DataPelago claims that its platform can increase performance tenfold and cut infrastructure costs by up to 80%, though these figures depend on each workload.
- The integration into NetApp products does not yet have a public timeline for availability.
DataPelago will operate as a subsidiary owned by NetApp. The acquiring company has not specified how many employees will join, the valuation of the startup, or the expected financial impact, so the immediate significance of the deal should primarily be measured by the technology acquired.
The move reflects a shift beginning to occur in the enterprise storage market. For years, the primary function of these systems was to store, protect, and serve data quickly. The expansion of AI adds another requirement: preparing large volumes of data before they can be used to train models, tune systems, perform semantic searches, or build applications with retrieval-augmented generation (RAG).
This work often involves extracting information from production storage, copying it to data lakes or processing clusters, and transforming it there. The process consumes time, network capacity, additional space, and computing resources. It also creates duplicate versions that must be protected, updated, and governed.
NetApp aims to reduce these movements by executing part of the processing at the layer where the information resides.
Nucleus brings CPU and GPU close to the data
DataPelago’s core technology is Nucleus, a processing engine designed to distribute operations across different hardware types. It can utilize CPUs, GPUs, and field-programmable gate arrays (FPGAs), depending on the characteristics of each task.
Its purpose is to perform transformations on structured, semi-structured, and unstructured data without forcing companies to transfer all data to an independent computing platform beforehand.
DataPelago presents Nucleus as a universal engine compatible with various analysis environments. The platform leverages open projects like Substrait and Apache Gluten to connect with query and distributed processing engines such as Apache Spark and Trino.
The company states that its technology can deliver performance up to ten times higher and reduce infrastructure costs by up to 80% compared to certain conventional approaches. These results are communicated by DataPelago and are not an automatic improvement applicable to all environments. The final outcome depends on data volume, query type, hardware availability, architecture, and integration costs.
The technical interest lies in bringing computation closer to storage. Moving petabytes of data for processing can be more expensive and slower than executing necessary operations where the data already resides.
This idea does not eliminate transfers entirely. Models, applications, and accelerators still need to receive information during execution. What it seeks to reduce is the systematic creation of intermediate copies used to prepare, filter, classify, fragment, or vectorize datasets.
What really means copy-free activation
NetApp uses the term zero-copy or copy-free activation to describe this approach. The phrase does not imply that no data will ever be copied, but rather that organizations could avoid complete and persistent duplications before starting certain AI and analytics workflows.
This distinction is important because enterprise data are rarely ready to feed directly into a model. They may be spread across files, databases, object repositories, documents, historical logs, and internal applications.
Before use, data typically need to be discovered, permissions verified, duplicates eliminated, governance policies applied, content extracted, fragments generated, and some information converted into vector representations. These phases can consume significant time and project budget.
NetApp aims to integrate DataPelago’s technology with its data platform, which includes ONTAP, disaggregated storage AFX, and AI Data Engine. The company already supported a model where AI tools approach the data rather than aggregating all data into a new platform first. This acquisition provides an accelerated execution engine to develop that concept.
However, NetApp cautions that these capabilities are part of its future technological direction. The company has not confirmed which products will incorporate Nucleus, when they will be available, or under what commercial model. Therefore, the acquisition does not mean these functions are already active in ONTAP, AFX, or AI Data Engine.
Storage wants to take on more roles in the AI era
The acquisition also illustrates how storage vendors are broadening their scope. The pressure no longer only comes from the need for more capacity or performance, but also from helping locate, protect, classify, and prepare information used by AI applications.
Companies have invested in GPUs, but these accelerators can remain underutilized when data arrives too slowly, are not properly prepared, or require multiple pre-transformations. This challenge has driven the development of architectures combining high-performance storage, fast networks, catalogs, vector engines, and governance tools.
With DataPelago, NetApp seeks to control a larger part of this journey. Instead of just delivering data to an external cluster, its infrastructure could perform some operations using CPUs and GPUs integrated within the data layer itself.
This move also aligns with AFX, NetApp’s disaggregated architecture for separating capacity, services, and control, and with AI Data Engine, which aims to discover and prepare enterprise data for AI projects.
The acquisition adds to NetApp’s collaborations with companies like Cisco, Google Cloud, Red Hat, and SK Telecom. However, DataPelago offers something different: proprietary technology and a team that the company can directly embed into its product strategy.
The key unknown is how NetApp will turn this engine into a manageable capability within heterogeneous enterprise environments. Performance is only part of the challenge. It also needs to address permissions management, workload isolation, format compatibility, observability, GPU utilization, and workflow management.
If the integration performs as promised, NetApp could reduce one of the most costly tasks in enterprise AI: moving and duplicating large datasets before they can be used. For now, the operation sets a clear technical direction, but its tangible benefits will be evident when DataPelago’s technology becomes available in customer products.
Frequently Asked Questions
What did NetApp buy?
NetApp has acquired DataPelago, a startup developing an engine to accelerate data processing with CPUs, GPUs, and other hardware types. The purchase price has not been disclosed.
What is DataPelago Nucleus?
Nucleus is a processing engine that distributes operations across different accelerators and aims to work on data near where it’s stored. It is focused on AI, analytics, and preparing large data sets.
What does zero-copy mean in this acquisition?
It means NetApp wants to reduce the need to create full copies of data on external platforms before processing. It does not mean that transfers will disappear entirely during operation.
Is DataPelago’s technology already in NetApp products?
NetApp has not yet announced an integration timeline or confirmed which products will feature these functions. The capabilities mentioned should be considered as part of their planned platform evolution.
Sources:
- NetApp, official announcement of the DataPelago acquisition, 07/16/2026.
- NetApp, technical analysis on Nucleus and copy-free data flows.
- DataPelago, technical information about Nucleus and heterogeneous processing.

