Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency

Qwen3.8-Flash-Next: A New Architecture, Towards Ultimate Cost-Efficiency

Qwen3.8-Flash-Next is an open-weights multimodal MoE model that previews the architecture for Qwen4. It introduces a hybrid Gated DeltaNet + Qwen Sparse Attention (QSA) for efficient memory and retrieval, a Gated Residual stream with four branches, N-gram Embedding to scale capacity cheaply, and the Muon optimizer. With 125B total parameters (6B active) plus 51B N-gram embeddings, it cuts training cost to about 1/9 of Qwen3.7-Plus while improving coding and office tasks. It natively supports 262K context, extendable to 1M, and the production version is priced at $0.16/M input and $0.47/M output tokens.

Put simply: GDN efficiently “remembers,” while QSA precisely “retrieves.”
  1. lnenad

    Adding to my homelab stack, hopefully it doesn't overthink like the little model. Actually, hoping it thinks a bit less. Wait actually I'm really praying it reasons a bit more directly. But wait, I'm really sure that it must be a bit better.

  2. hedgehog

    In initial testing on Ryzen 395 / Strix Halo it's about 22 tokens/s generation and the output quality is impressive. Better and faster than 3.8 27B, and enough better to justify moving away from 3.6 35B even though 35B is still faster. The Unsloth weights don't come with vision support but you can add that back yourself. llama.cpp recipe in case anyone else wants to save some time on setup: https://pastebin.com/fcqsbDTv

  3. rohansood15

    Didn't expect it to beat 3.8 27B so cleanly.

    Opus 4.6 Max self-hosted at 30 tok/s on a 5k Macbook in Aug 2026. The LLM timelines are crazy.

  4. respectattentio

    I can't imagine the future any more.

    US companies playing it safe and control models releases.

    Chinese companies are just like open source everything.

    It's like Chinese are incentivized to open source from day one (years ago).

    While most US companies are deciding in realtime.

    It's crazy that we need both to survive and advance further in the future we have never imagined.

  5. andy99

    > Qwen3.8-Flash-Next features a 125B-parameter main model, supplemented by an additional 51B N-gram embeddings, with 6B parameters activated per token.

    Didn’t see this mentioned yet. I wonder what this means for the effective size. It’s evidently ~176B paramètres, but how does that get quantized. A 4-bit quant under 100GB seems unlikely, I’m suspecting this won’t run in 128GB unified memory

    In principle I like the idea of trading more memory for compute though, even if there’s a memory shortage right now

  6. simonw

    I ran some pelicans at the four different reasoning levels (none, low, medium, xhigh - apparently high and xhigh are aliases of each other) on a DGX Spark using Unsloth's unsloth/Qwen3.8-Flash-Next-GGUF (UD-IQ1_S):

    https://tools.simonwillison.net/markdown-svg-renderer#url=ht...

    Surprised I didn't get one I liked as much as the Qwen 3.8 27B one https://simonwillison.net/2026/Aug/16/qwen-38-27b/#the-defau... , maybe because of quantization.

  7. a_humean

    Waiting for llama.cpp support to land, but this might be a big deal for Strix Halo users.

    6B active params helps around the memory bandwidth constraints, but a 128GB box can probably run the Q3/Q4 quants fairly easily with a decent context size. This might actually be better for strix users than 27B, which was already very good.

  8. tosh

    this is a new architecture (foreshadowing qwen 4)

    > trained at just 1/9 the cost of Qwen3.7-Plus, while outperforming it across the board

    https://x.com/Alibaba_Qwen/status/2092591393424515114

More from this day

2026-08-26