NetApp Acquires DataPelago for Storage-Layer AI Processing

NetApp acquired AI data infrastructure company DataPelago on July 16, adding its Nucleus engine for GPU- and CPU-accelerated processing at the storage layer. NetApp said Nucleus can process data in place rather than copying it to separate compute clusters, with claimed infrastructure-cost reductions of up to 80% and performance gains of up to 10x. The acquisition price was not disclosed.
NetApp acquired California-based AI data infrastructure company DataPelago on July 16, adding DataPelago's Nucleus universal data processing engine to its portfolio. NetApp did not disclose the transaction price, according to CRN.
NetApp said the acquisition adds GPU-accelerated data processing aligned directly with the storage layer. DataPelago's Nucleus uses heterogeneous CPU and GPU resources to process data where it is stored, rather than requiring data to be copied into separate analytics or AI compute clusters.
George Kurian, NetApp's CEO, said in the company's announcement: "As AI models and the chips that power them get ever more effective, enterprises need data infrastructure that is just as intelligent and powerful to harness the potential of their data." He added that DataPelago extends NetApp's ability to help customers understand and process data for AI workloads.
A zero-copy processing layer
According to NetApp, Nucleus can reduce infrastructure costs by up to 80% and deliver performance up to 10 times faster than conventional architectures. Those are company-reported, workload-dependent performance claims, rather than independently benchmarked results.
DataPelago's technology is intended to work between storage systems and data-processing frameworks. Blocks & Files reports that Nucleus uses the open-source projects Gluten, Velox, and Substrait, and integrates with engines and frameworks including Apache Spark, Trino, Flink, Ray, and Dask. The publication also reports integrations with SQL, Python, Airflow, Tableau, and Power BI.
CRN reports that DataPelago Accelerator for Spark can be deployed without rewriting applications or migrating data. That deployment model is relevant to teams with existing Spark pipelines whose data preparation stages are constrained by transfers between storage and GPU-backed compute environments.
Data management becomes an AI infrastructure concern
TechTarget characterized the acquisition as an expansion of NetApp's support for distributed data management, an area enterprises confront when preparing governed data for AI applications. The publication noted that NetApp had previously released data-management tools for collecting, curating, synchronizing, protecting, and preparing data for AI use, as well as an AFX all-flash array with disaggregated storage capabilities.
Rob Strechay, founder and principal at Smuget Consulting, told TechTarget that the technology could become a connective layer between file and object storage, enterprise data, and compute engines such as Spark. "By bringing GPU- and CPU-accelerated processing... directly to the storage layer, organizations should be able to execute analytics and AI data preparation without constantly copying data into separate compute clusters," Strechay said.
For data engineering and ML platform teams, the technical premise is familiar: movement, serialization, format conversion, and repeated copies can consume time and infrastructure budget before GPUs execute a training, feature-engineering, or inference-adjacent workload. Companies pursuing comparable storage-adjacent compute designs commonly seek to reduce those pipeline stages while maintaining governance controls around operational data.
The practical questions remain integration depth, hardware support, workload coverage, and whether reported gains hold across customer environments. NetApp declined to provide CRN with details beyond its press release and indicated that additional information would be forthcoming. StorageReview reports that DataPelago is operating as a wholly owned NetApp subsidiary following the transaction.
Key Points
- 1NetApp adds Nucleus to process AI and analytics data in place, targeting costly transfers between enterprise storage and accelerator clusters.
- 2Company-reported results cite up to 80% lower infrastructure cost and 10x faster processing, but performance remains workload and deployment dependent.
- 3Comparable storage-adjacent compute architectures can reduce data-pipeline overhead while raising integration, governance, and hardware-compatibility evaluation requirements.
Scoring Rationale
The acquisition is notable for enterprise AI and data-platform teams because it targets data movement and preparation bottlenecks around GPU workloads. Its relevance is strongest for NetApp customers and teams operating Spark, lakehouse, or distributed analytics pipelines, while transaction terms and acquisition integration details remain undisclosed.
Sources
Primary source and supporting public references used for this report.
View 4 more sources
- NetApp's DataPelago Buy Targets The Next Big AI Storage ...crn.com
- NetApp-DataPelago deal ups AI data management antetechtarget.com
- NetApp buys DataPelago to become full-stack AI data infrastructure providerblocksandfiles.com
- NetApp Buys DataPelago to Run GPU Data Processing Where the Data Already Livesstoragereview.com
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

