Samsung's LPDDR5X-PIM: 8x Memory Bandwidth, but a Software Nightmare

Samsung's Processing-in-Memory (PIM)

Samsung's LPDDR5X-PIM: 8x Memory Bandwidth, but a Software Nightmare

At Hot Chips 2026, Samsung detailed its Processing-in-Memory (PIM) implementation in LPDDR5X DRAM, placing MAC units inside each of the chip's 16 banks to exploit internal bandwidth. This yields 614 GB/s, eight times the standard 76.8 GB/s, and up to 2.4 TOPS per chip. While it works with standard memory controllers via special row addresses, the design forces software into a corner: PIM mode disables caching, prefetching, and out-of-order execution, and complicates multitasking. The article explores these challenges and suggests hardware changes to make in-memory compute practical.

PIM mode switching throws a wrench into the works for multitasking operating systems.
  1. bob1029

    The tradeoff with putting the compute in the memory is that you have to know exactly where the dependent information will be at all times. Most problems do not fit this pattern very well. AI, gaming and crypto being the most obvious exceptions. It is incredibly constraining to develop applications using specialized hardware like this. You might as well spin out an ASIC for whatever it is you are doing. All 3 applications noted above eventually got their own flavors.

    I think the Von Neumann bottleneck is mostly a feature. The fact that communication of information across distances is expensive should not be immediately assumed to mean that it is universally flawed to do this. You are paying for something when you use all those joules. I'd argue we are usually wasting our energy with regard to information communication (e.g., lighting up a network interface & copper because we couldn't be bothered to use SQLite), but other times this stuff is fundamentally required for practical solutions to exist.

  2. HarHarVeryFunny

    I remember taking VLSI design as part of my Comp. Sci. degree at Bristol, UK c.1980, using the Conway & Mead book, and "Commingling of Processing and Memory" was mentioned even back then.

    Obviously you (eventually) need your data where the compute is, especially in a non-von-Neumann architecture where moving data around isn't an option even if you were OK with the performance drop.

    It seems kinda obvious that eventually AI will be implemented as low power dataflow custom chips integrating memory/state & compute, but who knows!

  3. samuelknight

    I saw them present a similar concept at Hot Chips in 2020 or 2021. It's still a cool idea, however people should remember that there are like 20 of these exotic accelerators designs pitched at trade shows every year that go nowhere.

  4. londons_explore

    Whilst processing in memory is clearly the future, I am unconvinced by this implementation.

    Matrix multiplication involves getting every entry of the input and output matrices to be at the same multiplier at the same time. (Ie. N^2).

    To do that, a lot of data movement needs to happen. Movement is the main thing - the multiplication and addition is a sideshow as far as energy and silicon space is concerned. You need a 'around the chip' ring shift register to pass every element of one matrix past every element of the other.

  5. throwaway173738

    Might as well just go whole hog and change the entire computer architecture, then. A lot of the arguments against this change boil down to computers and software code don’t work well with this today.

  6. pragma_x

    What I find amusing about moving compute to a RAM bank is it _almost_ resembles where we were with ISA-based extended RAM back in the 1980's. Some cards featured a CPU that took over the whole system and/or functioned like an upgrade. Others were a "computer on a card" that provided other features. I think this goes to show how cyclic tech can be. So, something like Samsung's invention here might have gained traction, as overcoming the slow PC ISA bus would have been a huge accelerator, kind of like where we are now.

  7. consp

    So you basically dispose of cache for the memory region used? I wonder what the offsets of the cache misses is going to be in practice (the article addresses it but there is no solution/impact given by samsung).

  8. OptionX

    So instead of putting more cache on the cpu you just put the cpu on the cache.

More from this day

2026-08-29