Hot Chips 2026: High Bandwidth Flash Could Slash DRAM Costs for AI, but Software Hurdles Loom
Hot Chips 2026: Applying High Bandwidth Flash (HBF)

At Hot Chips 2026, a tutorial explored High Bandwidth Flash (HBF), a new memory technology that packages flash memory like HBM on the same chip as a processor. HBF offers far greater capacity than HBM but with lower bandwidth and block-based access, requiring software to treat it like a storage device. The talk examined how HBF could be applied to machine learning, including storing MoE experts and KV caches in flash, and replicating model weights to reduce cross-device communication. However, the software changes needed are substantial, and HBF's cost-effectiveness depends on workloads that aren't bandwidth-bound.
Taking advantage of SSD actually seems easier. The OS kernel can abstract away the difficulty of doing block-aligned accesses if you don’t use FILE_FLAG_NO_BUFFERING or O_DIRECT.
- rbanffy
I have a feeling Intel got rid of Octane a couple years too soon.
- xnx
Is this the same idea John Carmack had? (https://x.com/ID_AA_Carmack/status/2074248758422864226?lang=...)
"Memory cost and capacity are significant issues for AI accelerators.
Unlike game rendering, model inference can have a deterministic memory access pattern. You don’t need “random access memory” at all for model weights, and you could tolerate cold-start latencies in the multiple milliseconds, as long as continuous reads were delivered at the necessary bandwidth.
NAND flash is over 100 times cheaper per GB than HBM, so there should be opportunity there, even after giving a flash controller a 1024 bit interface with HBM bandwidth.
You could make a specialized pin protocol that just supported pipelined transfer of full 16KB+ pages from the flash to program-managed accelerator scratchpad memory and improve per-pin performance over HBM, but it might be more convenient to make it still look like a true random access memory with very fragile performance characteristics, where anything but sequential reads falls off a 1000x+ performance cliff.
That has the advantage of automatically using existing cache hierarchies, and providing a natural path to update the flash memory with new model weights. With the stream-to-scratch interface, code has to be completely rewritten before it works at all, while the ram-emulation interface will start off just extremely slow, and you can incrementally sort out the changes for full performance.
There may be cases where there isn’t enough scratchp […]
- rando1234
What will be the durability/lifetime properties of this technology in comparison to traditional DRAM?