THE APEX TIMES
NVIDIA at FMS pushes AI storage upgrades, open-sources cuFile to speed GPU access
As AI agents drive thousands of concurrent storage operations, NVIDIA says the bottleneck is shifting from GPU compute to the storage path. At this week’s Future of Memory and Storage conference, it unveiled new architectures and opened key APIs aimed at making secure, direct GPU-to-storage data access faster and easier to integrate.
NVIDIA used this week’s Future of Memory and Storage (FMS) conference in Santa Clara, California to argue that the next performance gains in AI will be constrained less by raw GPU power and more by the storage systems feeding “AI factories.” In a technical post focused on memory and storage pressures, the company said AI’s growing reliance on massive datasets and long context windows is pushing workloads beyond the practical limits of system memory, even when advanced memory is available.
The company’s central point is that increasing storage capacity alone does not solve the problem. Instead, storage has to be designed to handle surges in parallel requests, while continuously performing security and data-processing functions. NVIDIA said AI agents now consume very large amounts of data and, importantly, GPUs can directly initiate storage requests. That combination can generate thousands of concurrent operations hitting the same storage infrastructure.
NVIDIA described several functions that storage systems must perform at scale, including continuous encryption, compression, verification, and reconstruction of data. It warned that these “critical data services” can become bottlenecks when many agents access storage simultaneously. In other words, the storage path must sustain high throughput without falling behind the pace of accelerated compute, or the system’s end-to-end latency and data access efficiency suffer.
To make that case with numbers, the blog highlighted benchmark results involving NVIDIA’s Vera CPU, which is part of the NVIDIA Vera BlueField-4 STX. In a two-stage compression and encryption pipeline, NVIDIA said the Vera CPU can deliver up to 3.21 times higher throughput than an x86 CPU. The company tied that result to the idea that storage platforms can absorb high-volume AI data traffic more efficiently, potentially reducing the amount of compute infrastructure needed for the same storage-side processing load.
Storage integration also sits at the heart of NVIDIA’s software changes at FMS. NVIDIA said it is open sourcing its cuFile application programming interfaces (APIs) and the vertical storage software stack underneath them. cuFile is described as an open source component of NVIDIA GPUDirect Storage, a technology that lets GPUs, not just CPUs, read and write data directly to storage. NVIDIA said the approach uses hundreds of thousands of GPU threads along with other methods to achieve secure data access measured in microseconds.
NVIDIA also framed the cuFile open sourcing as a security-oriented interoperability effort. The company said it is helping unify a security-first storage stack based on Linux best practices, enabling interoperability between GPUs and data. It added that fast, secure access to data and storage can support preventive and detective cybersecurity measures, and that making cuFile openly available can help make security context, data, and storage accessible at the speed AI-driven defenses need. NVIDIA referenced the Open Secure AI Alliance as an initiative that aligns with that objective, and said the open APIs site is intended for contributions and cross-platform optimization with Google, Intel, NVIDIA, and Meta as inaugural maintainers.
Beyond software, NVIDIA announced an initiative called Storage-Next, aimed at aligning storage makers, controller vendors, thermal and cooling contributors, and orchestration operators on how GPU-driven storage should behave. NVIDIA said Storage-Next includes more than 40 leading storage and flash vendors, naming DDN, KIOXIA, and Micron among the participants. The initiative’s stated goal is to translate those behavioral alignments into interoperable, open industry standards, in part to support accelerated data access for large AI datasets.
Central to that effort is a framework NVIDIA called SCADA, short for scaled, accelerated data access. NVIDIA said SCADA allows massively parallel GPUs to pull only the data necessary for an application directly from storage into their own high-speed memory, rather than bringing broader data sets into memory in bulk. The post gave one integration example: DDN is said to be integrating SCADA with Infinia, DDN’s software-defined, AI-native data intelligence platform designed to eliminate storage bottlenecks at scale. DDN Chief Technology Officer Sven Oehme said the collaboration helps create a “more direct, efficient connection” between GPUs and data, keeping accelerated computing resources productive and speeding “time to insight,” a framing that ties infrastructure improvements to customer returns.
NVIDIA positioned these developments as part of its broader AI storage infrastructure work, including NVIDIA Vera BlueField-4 STX. It also described NVIDIA STX as a modular, rack-scale foundation powered by the NVIDIA Vera Rubin platform, NVIDIA Vera BlueField-4 storage processors, and NVIDIA Spectrum-X Ethernet networking. For enterprise security controls in the data path, the company said STX uses a unified NVIDIA DOCA security stack to enable continuous policy enforcement. It additionally mentioned NVIDIA CMX Context Memory Storage, which it described as an AI-native context tier built on STX for long-context, multi-turn, agentic inference.
Even with the detailed architecture explanations, some specifics remain absent from the blog. NVIDIA did not publish full benchmark methodology, workload profiles, or comparative results across a wider range of storage systems and AI agent behaviors. It also described its secure access approach using a split between “user parts” that remain outside the trusted computing base and a separate privileged component that configures protected access at setup, following standard Linux protocols, but it did not quantify operational overhead or compatibility constraints across different enterprise environments. Those details will likely matter to IT teams evaluating rollout timelines and integration risk.
What to watch next is how quickly cuFile and the related open source components are adopted by storage platforms and application developers, and whether Storage-Next’s standardization effort produces interoperable product updates that can reliably deliver the promised microsecond-class data access. Given the increasing emphasis on data path throughput, the near-term competitive battleground may shift from GPU benchmarks alone to end-to-end AI pipeline performance that includes encryption, compression, verification, and reconstruction under high concurrency.
Why It Matters
- If storage-side throughput and security processing remain bottlenecks, AI system performance could plateau even as GPU compute improves, making storage architecture a competitive factor.
- Open sourcing cuFile could reduce integration friction for enterprises and developers trying to enable direct GPU-to-storage access, potentially accelerating adoption of faster AI data pipelines.
- Storage-Next’s push toward common behavior and standards could influence how storage vendors and controllers design for GPU-driven workloads at scale.
- Benchmarks tying encryption and compression throughput to CPU choices suggest enterprises may need to reevaluate hardware balance across the data center stack, not just GPU selection.
Key Facts
- NVIDIA said AI workloads are increasingly limited by the storage path because AI agents generate thousands of concurrent storage operations, and system memory cannot hold the needed datasets and context windows.
- The company said storage systems must continuously encrypt, compress, verify, and reconstruct data, and these services can become bottlenecks when many agents access storage simultaneously.
- NVIDIA cited benchmarks showing its Vera CPU (part of NVIDIA Vera BlueField-4 STX) can deliver up to 3.21x higher throughput than an x86 CPU in a two-stage compression and encryption pipeline.
- At FMS, NVIDIA announced it is open sourcing cuFile APIs and the vertical storage software stack underneath them, describing cuFile as part of NVIDIA GPUDirect Storage that allows GPUs to read from and write to storage directly.
- NVIDIA unveiled Storage-Next, a multi-vendor initiative for aligning GPU-driven storage behavior and producing interoperable, open industry standards, and described SCADA to let GPUs pull only necessary data into high-speed memory.
- NVIDIA named DDN as integrating SCADA with Infinia and quoted DDN CTO Sven Oehme on improving the connection between GPUs and data to keep accelerated computing productive.
Technology Related
Palantir lifts after strong Q2 results and upbeat guidance, while markets track fresh Iran-Hormuz and Bessent comments
In late-morning trading, investors reacted to a widely cited market roundup that highlighted Palantir’s stronger-than-expected second-quarter performance and new guidance alongside ongoing macro headlines tied to the U.S. and Iran.
Palantir shares jump after market buzz, but details on the catalyst remain unclear
Palantir Technologies (PLTR) rose sharply on Aug. 4, according to a market write-up, without providing enough specifics in the available materials to pin down the exact driver.
Palantir shares rise on investor focus on AI embedded in its data platforms
A recent market write-up points to demand for Palantir’s AI capabilities, arguing they are being pulled deeper into the company’s core data offerings. The stock’s strong run over the past two years is tied to that shift, though the post provides limited granular detail on timing or customers.
Meta’s AI investment pitch boils down to ad payback, but the “upside” case depends on execution
A market analysis argues Meta’s biggest potential upside is relatively straightforward to model: the cost of building artificial intelligence systems should translate into higher value for advertisers, lifting what brands pay and supporting margins.
Darktrace and SECURE AI are among the first firms selected to feed risk outlines into Microsoft’s Agent 365
Microsoft says it is working with early cybersecurity partners, including Darktrace and its SECURE AI offering, to integrate “organization-specific” behavioral risk outlines into Agent 365. The goal is to help customers spot anomalous or suspicious agent activity that could indicate compromise.
NVIDIA backs NSF push for regional AI hubs to widen access to computing, data and training
The chipmaker says it will join a newly launched U.S. National Science Foundation initiative designed to bring shared AI infrastructure and education resources closer to universities and communities across the country.
Anthropic signs reported $10B, six-year computing deal tied to Nvidia-backed Volta Infra
The reported agreement underscores how major AI labs are locking in large-scale compute capacity and how Nvidia remains central to the supply chain for AI infrastructure.
Tesla vs. Amazon: Two AI bets, one spending-heavy robotics push and one profit-focused cloud build-out
A market piece from Yahoo Finance frames Tesla’s approach as cash-intensive robotics and robotaxi development, while portraying Amazon’s strategy as using its cloud infrastructure to convert AI demand into earnings.
Dinari, backed by Stripe and Apple alumni, teams with Circle to bring tokenized stocks to U.S. investors
A new tokenized-stock offering promoted through Circle will debut with a catalogue of 724 digital equities, including the full S&P 500, according to a report by Yahoo Finance.
Zacks points to AI push and AWS strength as Amazon’s growth drivers, while flagging mixed momentum across other stocks
An analyst blog cited Amazon’s AI-related initiatives, AWS performance and portfolio diversification as key supports for growth, in a broader market update that also highlighted strength in Marvell and Starbucks alongside notable challenges.