Qwen 3.8-Flash-Next: A 125B MoE Model Drops Tomorrow, Previewing Qwen4 Architecture

Qwen 3.8-Flash-Next releasing tomorrow (125B a6B)

Qwen 3.8-Flash-Next: A 125B MoE Model Drops Tomorrow, Previewing Qwen4 Architecture

Qwen team announces the upcoming release of Qwen3.8-Flash-Next, a multimodal mixture-of-experts model built on the next-generation Qwen4 architecture. The 125B-parameter model (with 6B active) will be available on ModelScope starting August 26, 2026, in both standard and FP8 versions. This early release aims to help the community prepare for the full Qwen4 family.

We are releasing these architectural advancements early to help the community prepare for the upcoming Qwen4 model family.
  1. SwellJoe

    Finally, a reason to own a 128GB Strix Halo or GB10 device. Or a reason to consider the new Mac Studio.

    I have a Strix Halo and dual 32GB GPUs in my desktop, and the latter is pretty much always better for running local models because it's quite a bit faster due to higher memory bandwidth. There simply haven't been any models that are better than Qwen 27B or Gemma 31B, which run comfortably in 64GB with big context.

    And, MoE should make it run at a close to usable speed.

  2. ddtaylor

    I enjoy the Qwen models a lot, but building things on top of them with OpenRouter has been painful.

    OpenRouter does a lot of great work and I really enjoy being able to use different models so easily. I like when a provider is phasing out an older model that still works for my needs and the price is much lower. It seems like such a good win-win.

    However, the problem is that many Qwen models have almost no capacity or is so flaky you literally have to just litter your code with a blacklist/whitelist of providers. OpenRouter has some attempts to solve this, but they don't work. In fact, OpenRouter has a lot of really cool stuff that is documented, but if you read the code it's not yet implemented or isn't actually there yet, which is a shame.

    I tried to get in contact with them at OpenRouter about this and I was interested in working with them in the past, but it's difficult to get in touch with the right people and they are growing very fast. I expect being acquired by Stripe will accelerate those problems in some ways. I have no doubt they will resolve all of these issues eventually and scaling that much that quickly is really hard, so kudos to them, but the road has been pretty lame and taken some wind out of my sails.

  3. notnullorvoid

    It will be interesting to see the intersection of this with inference engines like FreeToken which improve distribution of work for MoE models across CPU/RAM and GPU/VRAM.

    If all it takes for a competitive model to run locally at good speeds is a used 3090 and some DDR4, then we might be in for the year of local AI.

    https://github.com/FlashML-org/FreeToken

  4. fcanesin

    HF link: https://huggingface.co/Qwen/Qwen3.8-Flash-Next

  5. syntaxing

    Really looking forward to this, 27B is a struggle with a strix halo and Laguna 2.1 can do stupid things for tooling calls.

  6. vegnus

    I have an m1 max 64gb macbook. Anything I can do to get 3.8 27b at more than 10 tok/s or am I relegated? 3.6 a3b is good but its not as good

  7. big-chungus4

    > We are releasing these architectural improvements ahead of time so that the community can prepare for the upcoming full family of Qwen4 models.

    That gives me hope that "full family" means it will include smaller models like 4B.

  8. Catloafdev

    Very curious to see how this compares to Deepseek v4 Flash. I have to assume they wouldn't be releasing this if it was worse.

More from this day

2026-08-25