A field notebook on the platters, cells, and file systems holding up the AI build-out — vendors, capacity roadmaps, reliability data, and how infrastructure teams are actually tuning for it.
Every GPU cluster is now bottlenecked by something spinning or etched in silicon somewhere behind it. This issue maps the enterprise storage vendors and hyperscalers chasing the AI build-out, the hard drive and memory makers racing to keep up with it physically, what field failure data actually says about the drives carrying that load, and the handful of architecture decisions separating clusters that keep GPUs fed from ones that don't.
IDC's freshest tracker shows the fastest quarterly growth the external storage market has posted in years — and it isn't refresh cycles doing the pulling.
External enterprise storage systems revenue hit $9.2B in Q1 2026, up 22.7% year-over-year — the market's sharpest acceleration in recent memory, per IDC's Worldwide Quarterly Enterprise Storage Systems Tracker. IDC attributes the jump to AI infrastructure demand layered on top of deferred hardware refreshes and component price inflation working its way into system pricing.
Flash keeps winning the share fight, but note the last tile: all-HDD array revenue is still growing double digits. Spinning disk hasn't been pushed out of the AI stack — it's been pushed down it, into the cold and warm tiers behind the flash that's actually feeding GPUs. More on that split in section three.
The IDC top five by revenue, plus the AI-native challengers showing up in every DGX SuperPOD reference architecture whether or not they crack the revenue rankings yet.
By IDC's Q1 2026 revenue ranking, the external storage market is still led by the incumbents — but the products they're leading with are increasingly AI-specific rather than general-purpose:
Broadest portfolio in the market; leaning on an AI-storage attach strategy across PowerScale and PowerStore to ride server deals into storage revenue.
AFX disaggregated architecture and the new AI Data Engine (AIDE); DGX SuperPOD-certified, pitched on hybrid multicloud data mobility.
Formerly Pure Storage — renamed Feb 2026. Climbed to third on subscription-model adoption and AI-optimized platform revenue; FlashBlade still pushes past 10TB/s aggregate throughput with SLA-backed performance guarantees for AI pipelines.
Strong outside North America; among the fastest-growing suppliers by IDC's count over the past two trackers.
Rounds out the top five, folding storage into its broader AI factory and GreenLake positioning.
All-flash, disaggregated, deliberately un-tiered — the pitch is that tiering itself is the bottleneck it removes.
WEKApod appliances claim 17M IOPS and 800Gbps in a compact footprint, built around DGX SuperPOD integration.
Exabyte-scale S3-compatible object storage with native GPUDirect support at 200GB/s straight to the GPU.
Positioned squarely around "AI factories" — new performance, efficiency, and security tooling unveiled at ISC 2026.
Holds roughly 17% of the parallel file system market; the Storage Scale System 6000 is NVIDIA-certified for metadata-heavy training jobs.
Named a Leader in the 2026 GigaOm Radar for object storage, cited for optimization and enterprise-scale strength.
"AI-native" here means a vendor built its architecture around GPU-adjacent throughput from the outset, rather than adapting an existing enterprise array line — it isn't a claim about revenue rank.
AWS, Google Cloud, Azure, and Oracle have each split storage into GPU-throughput and high-TPS tiers sold as distinct products — the same split showing up among the AI-native challengers in section two.
FSx for Lustre pushes 1TB/s aggregate throughput with GPUDirect Storage straight to GPU memory; S3 Express One Zone adds a single-digit-millisecond, high-TPS object tier built for checkpoint-heavy training. Meta FAIR has reported sustaining 140 Tbps and 1M+ transactions/second across 60PB of co-located training data on it.
Managed Lustre delivers up to 10TB/s over RDMA to its largest TPU and GPU training clusters. Rapid Storage throughput jumped from 6 to 15TB/s this year, and a new Smart Storage layer adds semantic search over unstructured files for agent workloads.
Azure Managed Lustre (AMLFS) is in preview at 25PiB namespaces and 512GB/s. Blob Storage remains the default for training and fine-tuning datasets at enterprise scale, with Elastic SAN block storage rebuilt for agents that query roughly an order of magnitude more than typical users.
Oracle's AI storage story runs through Supercluster networking rather than a dedicated storage product line — Acceleron RoCE claims roughly 2x the storage IOPS of its prior generation, feeding superclusters Oracle says can scale to 800,000 GPUs.
The pattern across all four is the same one running through the enterprise vendor field: nobody's shipping one array and calling it done. Each cloud now sells a bandwidth-first tier for feeding accelerators during training and a latency/TPS-first tier for checkpointing and high-frequency object access — priced, and marketed, as separate decisions.
Cold and warm tiers behind the flash still run on spinning disk — and the two companies that make almost all of it are both about to switch recording technologies at once.
Western Digital and Seagate currently top out around 26–32TB per drive using conventional and shingled recording. Both are moving to heat-assisted magnetic recording (HAMR) to get past that ceiling — a genuine technology transition, not just a density bump, and both companies are racing to be first through it cleanly.
| Vendor | Shipping now | Next milestone | Timeline |
|---|---|---|---|
| Western Digital | 26TB CMR / 32TB UltraSMR | 36TB CMR / 44TB UltraSMR (HAMR) | Debut late 2026 · volume H1 2027 |
| Seagate | Mozaic 3+ platform (HAMR, shipping) | 50TB (record-setting) | Qualification target 2027 |
| Toshiba | Up to 20TB-class, CMR-focused | — (limited public roadmap) | — |
Western Digital's longer-range target is 80TB CMR and 100TB UltraSMR by 2030, stacking HAMR with its OptiNAND caching layer and mechanical refinements. Seagate is publicly betting it can move faster through the HAMR transition than WD did — both companies spent years on now-abandoned interim technologies (WD on MAMR) before converging on HAMR as the only real path past today's ceiling.
Why it matters for AI infrastructure specifically: checkpoint and dataset archives are growing faster than GPU budgets, and the cost-per-TB of the cold tier is what keeps the flash tier affordable. A stalled HAMR ramp at either vendor would show up as pricing pressure across the whole stack within a couple of quarters.
HBM is the scarce input training and inference clusters actually fight over. NAND is riding the same wave a step behind, on the back of a spinoff few saw coming.
| Vendor | HBM position | Notable figure |
|---|---|---|
| SK Hynix | Market leader | ~58–62% HBM share · ~72% operating margin |
| Samsung | Closing the gap | ~21% HBM share · record ~₩57.2T quarterly profit |
| Micron | Behind on HBM timeline | DRAM revenue +81.6% QoQ |
The bigger surprise this year is on the NAND side. Kioxia and SanDisk — split off from Toshiba and Western Digital, respectively — have been the standout gainers as data center NAND demand and their post-spinoff valuation re-rating compounded together. Kioxia's quarterly operating profit jumped roughly 15-fold to ¥596.8B, briefly making it Japan's most valuable listed company by some measures.
Backblaze's Q1 2026 Drive Stats report remains the largest publicly published dataset on real-world HDD failure — over 300,000 drives, tracked quarter over quarter for 13 years running.
The headline paradox is worth sitting with: the fleet's quarterly AFR ticked up versus the prior quarter, yet the newest, highest-capacity drives are the most reliable cohort in the dataset. Higher areal density hasn't come at the cost of reliability — if anything the opposite, at least so far.
| Model | Capacity | Quarterly AFR |
|---|---|---|
| WD WUH722222ALE6L4 | 22TB | 0.38% |
| Toshiba 20TB-class | 20TB | 0.92% |
| HGST HUH721212ALE604 | 12TB | 2.65% |
| Seagate ST12000NM0008 | 12TB | 2.83% |
| HGST HUH721212ALN604 | 12TB | 3.98% |
| Seagate ST10000NM0086 | 10TB | 4.63% |
| Seagate ST14000NM0138 | 14TB | 4.90% |
Backblaze's own caveat is worth repeating rather than smoothing over: several of the higher-AFR models above have small drive populations left in the fleet, and low sample sizes make quarterly AFR swing hard on individual failures. Treat single-quarter spikes on low-population models as noisy; treat the 20TB+ pooled figure, backed by a much larger sample, as the more durable signal.
Five decisions that show up repeatedly in how infrastructure teams keep storage from starving expensive GPUs.
Moving storage traffic to NVMe over Fabrics cuts latency from milliseconds to tens of microseconds. Paired with GPUDirect Storage it's been measured at 351 GiB/s sequential reads, and can lift GPU utilization 2–3x in clusters that were I/O-bound rather than compute-bound.
Routes data straight from storage into GPU memory, skipping the CPU entirely — 40+ GB/s direct transfer in practice. It matters most for checkpointing, where the write has to land inside a fixed time window without stalling the training step behind it.
Pick by access pattern, not brand recognition: Lustre (~41% of this market) leads on sustained sequential bandwidth; IBM Storage Scale (~17%) handles metadata-heavy jobs with many small files; WekaFS (~6%) is NVMe-native and built for mixed, unpredictable I/O rather than disk-era sequential patterns.
The pattern that's converged across large training runs: a fast tier writes to node-local NVMe every few minutes, a mid tier propagates to shared storage roughly every 30 minutes, and a durable tier lands in object storage every few hours — draining asynchronously so training never blocks on the slowest tier.
The recurring mistake is buying one architecture for every job. Sequential training I/O wants Lustre-class bandwidth; mixed inference and agentic workloads want WEKA- or VAST-style flexibility; anything running on NVIDIA reference architectures needs GPUDirect-certified storage from day one, not bolted on after the first bottleneck shows up.