Technical Blog

DDN AI400X3M Doubles Read IOPS Without Adding a Watt

Enterprise AI teams, NVIDIA Cloud Partners (NCPs), and Sovereign AI programs are running into the same constraints. Training teams need to complete more work inside a fixed cluster budget. Inference teams need more tokens per watt, more tokens per dollar, and better answers from agentic and long-context workloads. GPU capacity, power, rack space, and capital are not expanding as fast as demand for AI training and inference. When a training run or inference request waits on storage, the GPU continues to consume power without advancing the work.  

The newly available EXAScaler 6.3.9 software release gives operators per-tenant I/O dashboards, volume-capacity enforcement, and client-side encryption for shared storage. More importantly, the DDN AI400X3M appliance combines that software with a hardware platform delivering 190 GB/s of sequential read throughput, 110 GB/s of sequential write throughput, and 8M peak read input/output operations per second (IOPS) in two rack units (2RU) at 2.5 kilowatts (kW). 

The Same 2.5 kW Now Does More AI Storage Work

The AI400X3M keeps the 2RU, 2.5 kW appliance envelope of the previous AI400X3I appliance while raising sequential read throughput from 140 GB/s to 190 GB/s and peak read IOPS from 4M to 8M. Sequential write throughput remains 110 GB/s. At constant power, sequential reads rise from 56 GB/s per kW to 76 GB/s per kW, while peak read IOPS rise from 1.6M per kW to 3.2M per kW. That is 36% more sequential read throughput and twice the peak read IOPS within the same appliance power rating. 

By reducing data stalls, a faster storage path makes more of each GPU-hour available to the workload and improves one of the inputs to tokens per watt and tokens per dollar. The result also depends on the model, serving stack, network, cache behavior, production request mix, and output quality. 

Recent DDN work on inference efficiency defines inference efficiency as production spend divided by successful outputs that satisfy the user. Answer quality remains part of the calculation because rerunning output that fails the quality gate consumes more GPU time and power. 

The DDN work on Golden Traces shows why request mix matters when an inference test is meant to represent production. Average the requests together, and the pileup disappears on paper. In production, several cache misses can land at once, drive up storage traffic, and leave GPUs waiting for data. Those GPUs keep drawing power while the job waits, pushing up the GPU-hours behind every successful answer. 

AI400X3M gives the storage path more headroom to prevent potential pileup. The chart below shows a comparison inside the rack and power envelope reported for both appliances. 

Figure 1. AI400X3M raises sequential read throughput 36% and doubles peak read IOPS inside the same appliance envelope. 
The 2RU, 2.5 kW unit delivers 190 GB/s of sequential read throughput and 8M peak read IOPS. 

Shared Capacity Needs a Tenant Boundary

An NCP has to keep one customer workload from consuming the storage capacity allocated to another. Enterprise AI teams face the same problem across business units, while Sovereign AI programs may separate agencies, research groups, or regulated workloads. Pooling improves utilization only when each tenant keeps its allocation and data boundary. 

EXAScaler 6.3.9 puts the controls in the storage path. Per-tenant I/O dashboards in Grafana show consumption and performance for each tenant. Volume-capacity enforcement gives operators a boundary they can apply. Client-side encryption protects the data, while isolation across compute, network, and storage separates each tenant. 

Per-tenant dashboards give operations teams the evidence to find the tenant creating pressure, enforce the allocation, and decide where the next unit of capacity belongs. Billing remains a separate service function. 

Hot Pools observability adds per-storage target statistics. When data moves between flash and disk, the team can see which tier is active instead of inferring it from a slow job. 

EXAScaler 6.3.9 also moves alerts to AlertManager and VictoriaMetrics. Configuration snapshots capture the pre-upgrade state. Partial resume, failure classification, and guards on irreversible controller steps show the operations team where an upgrade stopped and which steps can continue. 

Keeping customers or business units in one shared system avoids a separate capacity island for every tenant, while identity, capacity, and encryption controls preserve the boundary. The diagram below tracks one tenant identity across the storage controls and the operator view. 

Figure 2. One tenant ID connects storage enforcement with the operator view. 
The request reaches a shared namespace through isolation, volume-capacity, and encryption controls. Grafana reports storage I/O consumption and performance for the same tenant. 

Partners Need a Repeatable Unit for the Next Rack

To repeat the design, a partner or internal platform team needs a fixed answer for what to order, how it fits into the rack, and what the support team will see after deployment. 

AI400X3M packages those decisions as a turnkey DDN AI appliance. A team chooses combined data plus metadata or dedicated data and metadata, network links from 200 to 800 gigabits per second, and a 120 TB, 250 TB, or 500 TB usable capacity bundle built with Non-Volatile Memory Express (NVMe) drives. 

For a standard deployment, those bounded choices replace a custom bill of materials. The starting points are specific: NVIDIA compute, high-performance networking, AI400X3M or ES400NVX2 storage, an ES400NVX2 management node, and the DDN management stack. Partners can deploy a pre-validated NVIDIA GB200 NVL72 configuration or size another supported NVIDIA compute and network configuration. 

For an NCP, known units remove some of the variability from infrastructure planning. The service model still needs workload-specific token throughput, power, quality, and GPU-hour cost. Putting those inputs against the fixed appliance envelope shows how many racks fit the customer’s demand and what each rack must return. 

One combined AI400X3M provides 190 GB/s read, 110 GB/s write, and 8M peak read IOPS in 2RU at 2.5 kW. Up to 18 combined appliances can operate in one namespace at the stated upper test limit. Eighteen units produce 3.42 TB/s read, 1.98 TB/s write, 144M read IOPS, 9 petabytes (PB) of usable capacity, 36RU, and 45 kW. Every fleet number is the one-appliance number multiplied by 18. 

Figure 3. One AI400X3M becomes the sizing unit for an 18-appliance combined namespace. 
The exact multiply reaches 9 PB usable capacity, 3.42 TB/s read, 1.98 TB/s write, and 144M read IOPS in 36RU at 45 kW. 

AI400X3M sets the storage-side working point. A workload replay measures how much of that gain reaches the training or inference job. 

For training, record checkpoint time, recovery time, GPU utilization, and wall power. For inference, record tokens per hour, storage wait time, wall power, and the quality-gate pass rate. Attach the fleet GPU-hour price to either run. 

The resulting data connects the reported 190 GB/s and 8M peak read IOPS to productive GPU time, tokens per watt, tokens per dollar, and service margin.  

Bring the Workload to Fully Connected 2026 or SC26

Meet DDN at CoreWeave Fully Connected 2026, September 29 through October 1 in San Francisco, or SC26, November 15 through 20 in Chicago. Bring one production trace, the rack power limit, and the fleet GPU-hour price. We will map the workload to AI400X3M and define the replay. 

Explore these additional resources to find out more! 

Explore our Resources
September 29 - October 1, 2026San Francisco, CA
CoreWeave Fully-Connected | San Francisco, CA
Events