Weekly signal for the people building and buying AI infrastructure — what moved, why it matters, what's next.
A quick-turn issue this week: a physical limit in HBM packaging just pushed a roadmap milestone out a full generation, Gartner published its enterprise storage rankings, and two NAND makers put a number on how seriously they're taking the capacity crunch. Four stories, one deep dive, one term worth knowing.
Six things that happened this week, and why I'd pay attention to each one.
Same six Leaders as last year — Everpure, Huawei, HPE, NetApp, Dell, IBM — but NetApp slipped from second place to behind Huawei and HPE. Everpure leads both axes for the second straight year.
Roughly ¥5 trillion earmarked to expand Japanese NAND output, as flash capacity keeps getting redirected toward AI-driven demand.
An incremental refresh to IBM's enterprise flash array line, aimed at keeping DS8000 competitive on raw throughput.
Nutanix reported continued growth even as the underlying hardware it bundles gets more expensive — a direct downstream effect of the flash and DRAM price run-up.
The Abaco Project is a rack-scale system with over 100TB of CXL-attached shared and pooled memory, built for Pacific Northwest National Laboratory's next-generation AI and HPC workloads.
Rubrik's fiscal Q2 2027 revenue came in well ahead of its own guidance, continuing a run of beats.
The Storage and Memory Signal reaches people who actually buy and operate AI storage and memory infrastructure, AI DevOps engineers, infra leads, and the executives they report to. If that's your buyer, this slot is available.
A story instead of six headlines, since this one changes a roadmap most of the industry was counting on.
At Hot Chips 2026 on August 23, SK hynix said what a lot of people in the HBM supply chain had been quietly worrying about: hybrid bonding won't be ready in time for HBM4E. The industry's most anticipated memory-packaging transition is being pushed out to HBM5 at the earliest, which means the current MR-MUF (mass reflow molded underfill) packaging process is going to have to stretch further than planned — including through Nvidia's Rubin generation.
The 775-micron figure isn't arbitrary — it's set to match the thickness of a standard 300mm logic wafer, so the HBM stack sitting next to the compute die doesn't create a step height that breaks the packaging process. As HBM stacks add more DRAM layers, each layer eats into that fixed budget. Hybrid bonding — which fuses dies directly, copper-to-copper, without the solder bumps traditional packaging relies on — was the planned way to keep adding layers without blowing through the ceiling. SK hynix's admission that it isn't ready in time means the industry has to keep stretching MR-MUF, a solder-based approach that's been the workhorse packaging method for HBM3 and HBM3E, further than most roadmaps assumed it would need to go.
The practical effect: HBM4E-generation accelerators, including designs slotted for Nvidia's Rubin platform, will ship on packaging technology that was supposed to have been retired by then. That's not necessarily a capacity problem in the near term, but it does mean the industry is now carrying a known technical debt into HBM4E that hybrid bonding was meant to pay off. Whoever gets hybrid bonding production-ready first for HBM5 gets a real packaging advantage — this is now a race worth tracking the way HBM capacity itself has been tracked all year.
One piece of storage or memory vocabulary, explained properly, every week.
In plain English: it's a way to let multiple servers share one big pool of memory over a fast interconnect, instead of each server being stuck with only the DRAM physically plugged into it.
Compute Express Link (CXL) is an interconnect standard that runs over the same physical connector as PCIe but adds a cache-coherent protocol on top, so a CPU can treat memory sitting outside its own motherboard almost like local DRAM. "Pooling" is the specific trick of taking a bank of CXL-attached memory modules and letting several hosts draw from it dynamically, rather than statically wiring a fixed chunk to each server. If one server's workload needs more memory this hour and another's needs less, the pool can shift capacity between them instead of leaving DRAM stranded on an idle machine.
That matters for AI infrastructure specifically because GPU memory is expensive and workloads are lumpy — training jobs, KV cache demand, and batch inference all spike memory needs at different times. A fixed per-server memory allocation means you provision for peak and eat the idle cost the rest of the time. Pooled CXL memory lets a cluster carry less total DRAM for the same effective capacity, because it's shared rather than siloed.
This week's Micron/Primemas Abaco system for PNNL is pooling at an extreme end of the spectrum — over 100TB in one rack, built for a national lab's HPC and AI workloads. Most enterprises won't run anything close to that scale, but the same underlying mechanism is what CXL vendors are increasingly pitching for ordinary GPU clusters: less wasted memory, more flexibility, without redesigning the compute side.
AI Infra Summit 2026, Santa Clara. About as close as this space gets to a dedicated conference. Expect vendors to time announcements around it.
Q3 earnings season for storage and memory vendors (Samsung, SK Hynix, Western Digital, Seagate). Watch for capex guidance updates and anything on HBM4/HBM5 qualification, given this week's packaging news.
SC26, Chicago. The HPC/storage world's biggest annual gathering. Parallel file system and interconnect announcements tend to cluster here.