Technical Blog

DDN Infinia 2.4: The Foundation for Higher GPU Utilization in Production AI Factories 

AI has crossed a threshold. The question is no longer whether AI works. It is whether AI pays. Boards are asking about cost per token, GPU utilization, and time to revenue in production. AI infrastructure is now expected to be an economic engine, not an experiment. 

AI Training rewarded raw bandwidth. Production AI inference puts many tenants on the same data at once, securely and without pause. The data layer either keeps every GPU earning or becomes the reason it does not. 

DDN Infinia 2.4, announced at the RAISE Summit in Paris, delivers the enterprise foundation for running AI factories in production: multi-tenancy, identity and security, governance, and expanded protocol support, on the architecture already powering the largest AI deployments in the world. 

GPU Utilization Is a Data Problem Before It Is a Compute Problem 

Most organizations have solved the GPU acquisition problem. They have them running, but what they have not solved is the return on them. GPUs sit idle waiting on data. Inference costs climb as agentic workloads multiply the compute and data movement behind every interaction. Security reviews stall deployments for months. Teams that want to share expensive infrastructure stand up separate clusters because their data layer cannot isolate tenants safely. 

Infinia addresses each of those where it starts. Customers running DDN Infinia reach time-to-first-byte 25x faster, cut cost per token by 75%, and drive GPU utilization as high as 99%. When the data layer keeps GPUs fed, the economics of the entire AI investment change. 

What Infinia 2.4 Delivers 

  • Production-grade multi-tenancy. Run many teams, business units, or customers on one platform, with per-tenant isolation, quota enforcement, and governance controls. For NVIDIA Cloud Partners (NCPs) and managed AI operators, this is the difference between a cluster and a business: billable, isolated capacity – without a separate environment per tenant. 
  • Enterprise identity and security. Infinia 2.4 plugs into the identity stack enterprises already run, with integrated identity management and fine-grained access control. AI initiatives stop stalling at the security review, with regulated and sovereign environments getting the governance their mandates require. 
  • Native POSIX support. Even the most object-native AI environment touches files daily: Tools that expect a path, scripts that write checkpoints, preparation steps built on file utilities. Those moments no longer require a separate storage system, Infinia has a native POSIX client, not just a gateway. And when file performance is a critical need, at training scale and beyond, EXAScaler is the engine for that job. 
  • S3 compatibility and continuity. Existing S3 applications, SDKs, and workflows carry forward unchanged, so teams gain the performance and scale of Infinia without rewriting what works. 

Infinia also advances the economics of object-native inference: sub-millisecond access to datasets and model artifacts, massive concurrency for multi-tenant inference, Retrieval Augmented Generation (RAG) with time to first token 22x faster than AWS Express, and distributed KV cache acceleration over RDMA direct to GPU memory, delivering up to 4.2× faster time-to-first-token. As inference becomes the dominant cost in AI operations, these capabilities decide whether an AI factory runs profitably. 

Why Architecture Decides Sustained GPU Acceleration 

Feature checklists change every quarter. Architecture does not. When evaluating an AI data platform, the questions asked should look at the architectural core of the platform.  

  • Was the platform designed for AI or adapted to it? Infinia log-structured architecture and per IO data protection were built for AI access patterns. The difference is invisible in a demo and clear in production, which is why performance holds steady as tenants, protocols, and workloads multiply. 
  • Does every request pay an architectural tax? Some designs consult a separate metadata service before every fetch. Some fix data protection at cluster configuration rather than deciding it per IO. Each choice costs almost nothing at demo scale, but compounds at production scale, exactly when miss bursts, mixed workloads, and tenant concurrency arrive together. 
  • Can performance and capacity scale independently, with tenancy enforced in the data path? Infinia disaggregated design scales each dimension on its own terms, so customers grow into demand instead of overprovisioning ahead of it. 

A Data Layer That Strengthens the Stack You Already Run 

One more architectural choice shapes enterprise AI: where the data platform stops. Enterprises have spent years building their analytics estates on platforms like Snowflake and Databricks. When a data platform expands upward into query and analytics that replace these well-known environments, it forces a choice between the infrastructure and the tools your teams already depend on. 

DDN takes the position that Infinia is the high-performance data foundation beneath your AI and analytics stack, built to make the platforms you have chosen faster, more governed, and more economical, not to replace them. Your teams keep their tools, and your AI factory gets a data layer that serves the whole pipeline, from data preparation through training, inference, and agentic memory. 

One Platform, Two Engines 

Infinia is one of two engines delivering the DDN Data Intelligence Platform, alongside EXAScaler, the parallel file system behind many of the most demanding AI and HPC environments on Earth. Modern AI pipelines are multi-protocol: preparation and training frequently speak file, while data lakes, inference services, and agentic memory speak object. 

What separates the engines is where a workload center of gravity sits. When the environment is object-first with file at its edges, Infinia powers it, and the new POSIX support exists for exactly those edges. When file performance decides outcomes, at training scale and beyond, EXAScaler powers it. Customers get maximum performance on both from one partner. 

Proven Where It Counts 

Organizations including NVIDIA, xAI, Salesforce, Mistral, SK Telecom, TotalEnergies, and leading government and research institutions rely on DDN to power some of the largest AI environments on Earth. They chose DDN for the reasons Infinia 2.4 matters now: higher utilization of every GPU dollar, and faster paths from model to production. Explore Infinia at ddn.com/Infinia.  

What is GPU utilization?

GPU utilization is the percentage of available GPU compute actively doing work rather than waiting. In AI environments, low utilization usually signals a data delivery problem, not a compute shortage.

How does storage affect GPU utilization?

Storage sets the pace at which data reaches the GPU. When the data layer cannot sustain throughput at low latency, GPUs stall between batches. GPU direct storage paths that move data straight into GPU memory remove that wait, and customers running DDN Infinia drive GPU utilization as high as 99%.

What is new in DDN Infinia 2.4?

Infinia 2.4 adds production-grade multi-tenancy, integrated enterprise identity management, fine-grained access control, and POSIX support, alongside continued S3 compatibility.

How does multi-tenancy help AI service providers?

Per-tenant isolation, quota management, and governance controls let providers sell billable, isolated capacity on shared infrastructure instead of building a separate cluster for every customer.

Explore our Resources
October 21, 2026New York, NY
STAC Summit 2026 Fall | New York, NY
Events
August 31 - September 3, 2026Riyadh, SA
LEAP 2026 | Riyadh, SA
Events