What happens when a GPU reads memory

What happens when a GPU reads memory

A deep dive into the hardware path of a single GPU memory read, following a vector-add kernel's global load from SASS instruction to DRAM. The author reverse-engineers the RTX 4090's memory system, revealing the costs of L1 hits, TLB misses, L2 hits, and DRAM accesses, and explaining the roles of the coalescer, L1 cache, TLB, crossbar, L2 slices, and memory controller.

One LDG.E asks for four bytes in each of 32 lanes; serving it takes four 32-byte sectors, one cache line, one address translation, a crossbar crossing, one of thirty-six L2 slices, and, when it misses everywhere, an activate and four column reads at a DRAM chip.
  1. empiricus

    For a long time, the chip manufacturers had an inclination to simplify the hardware and rely on the software adapting and optimizing. But for decades this bid failed. Now we have the unrelenting AI capable of finetuning kernels relatively quickly. Maybe simpler hw will work this time? Note: not sure if TPU/NPU is not only simple but also too limited.

  2. xyzsparetimexyz

    > Little of the detail of this path is documented by NVIDIA, at least not to the level that we’d like, so we’ll determine it by running timing experiments on the hardware itself

    Or you could just use the AMD isa.

  3. KellyCriterion

    VERY good article, reminds me on:

    "what every programmer should know about memory"

    https://github.com/Ty-Chen/Reading-List/blob/master/What%20e...

More from this day

2026-08-21