Slotstream - Run 104GB Qwen3.8-Flash-Next on any Mac by streaming experts from SSD
Show HN: Running 104GB Qwen3.8-Flash-Next on 48GB Mac with at ~12 tok/s
Slotstream lets you run the massive Qwen3.8-Flash-Next model (125B MoE, 104GB at 4-bit) on Macs with far less RAM by streaming experts from SSD. It uses MLX and Swift, and provides an Ollama-compatible API. On a 48GB Mac, it achieves ~12 tok/s warm decode, with cold start to first token in ~3s, and peak memory of 32GB (auto-sized, cappable). The tool requires ~110GB free disk, works on Apple Silicon with macOS 14+, and includes a doctor command to check compatibility before downloading. It supports Ollama and OpenAI chat/generate endpoints, streaming, and various sampling options. Memory is elastic, resizing between requests. Installation is a simple curl script, and the binary is small; the weights are a one-time 104GB download. Slotstream is ideal for developers and AI enthusiasts who want to run large models locally on modest hardware.
Cache size changes speed, never output. Greedy decoding is byte-identical between a 4 GB cache and a 24 GB one, and that equivalence is a standing test.
- embedding-shape
> Hugging Face is the bottleneck, not your link.
README could clearly make use of a cleanup, seems to be more like a session log dump now than a good introduction to the project for a new user. Maybe try something like "Remove anything from the README.md that wouldn't be helpful to someone who sees this project with zero context, for the first time. Rewrite all paragraphs and sections to be concise and remove all fluff, leave only important details new users must know before using the project".
- mulemisterX
I have a 48GB M5. I don't need to run larger models. I want more context. I've managed to set the context window at 71,680 using Qwen3.8-27B-oQ4e-fp16-mtp. But I want more. Is anybody, with similar specs, able to set their context window higher?
- prometheus1992
It's hard to believe 16GB unified memory will give you 5 tok/sec unless you are ignoring the thermal warnings. I am running Qwen3.6-35B-A3B on my 16GB M3 and get 7-8 tokens/sec with all the optimizations while keeping the peak memory and thermal warnings at check. https://github.com/deepanwadhwa/samosa-chat
- whartung
I'm hoping to see progress in this space.
Folks talking about how 32G is not enough for local use, but then there's been work like this to empower it.
My hope is that the new 32G M6 will be "useful" locally, possibly because of work like this.
- jwr
For what it's worth, I've spent some time with Claude to develop a local runner for `llama.cpp`. I run Qwen3.6-35B-A3B-MTP (fast!) and Qwen3.8-27B (20 tokens/s).
This was definitely worth the effort. Benchmarking and testing various approaches and various options really paid off. For example, one thing that surprised me was that MTP made things slower, not faster for Qwen3.8-27B.
I use a 64GB MacBook Pro (M4 Max).
- jmward01
Not a mac/UMA discussion point, but is it time to add additional, installable, DDR5 to GPUs? I can see this as a win/loose. PCIe 5x16 is close to maxing out the bandwidth available from high end dual channel DDR5 now, but not quite. I'm not a hardware person but I suspect putting it on the card could lead to significant performance improvements over using system ram so allowing systems like this, where MOE weights are shed, to get even higher performance than just adding that DDR5 to the system. Bigger models become closer to reality and it provides more of a pathway for developing technologies that take advantage of it. Of course the loose side is that you just put a lot of specialized ram on a card instead of into the system where it could be used for other things. I could see a place for a 16GB card with 64GB(or more) of DDR5 especially if we start seeing MOE and similar technologies really start being designed for this concept.
- jacquesm
I love these efforts to get proper models running on lower cost hardware and I think this is where the next real breakthrough will come from. The more efficient this sort of thing can be done the bigger the chance to democratize this tech, 'good enough' is what you need and as long 'top of the line' gives a competitive edge even if it is at a cost there is a substantial risk of the door closing on general computing at some point in the near future. Keep in mind that there is no guarantee that the pendulum has to swing back, it can swing one way and get stuck, and then you're going to have to beg for crumbs from the haves.
- atif089
As someone who is just looking at the theoretical benchmarks of each of these models I'm curious if anyone could share what are the problems (maybe around code) that flash-next was able to solve which 27b was not able to