Xiaomi's Mimo 2.6 RL Dashboard Shows 1,192 Accepted Samples in One Minute

Xiami Mimo 2.6 Live Training Dashboard

Xiaomi's live training dashboard for Mimo 2.6 reveals the raw mechanics of reinforcement learning: accepted samples, pass rates, and reward signals update in real time. At one point, 1,192 of 1,568 samples were accepted with a pass rate of 0.664 across 7,826 judgments. The dashboard offers a rare, unfiltered look at how a large model's RL training actually progresses.

step 15 done · dynsam/avg@n 0.596 ▲0.012 2.69B tok
  1. joelwallis

    I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve.

    The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.

    --

    PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.

  2. dr_dshiv

    Well, if open source AI is dangerous (for OpenAI/Anthropic IPOs?), this is like watching a time bomb.

  3. passive

    Neat! I've been trying out their next model for the last week, which I assume is a version of this, and it's been a good experience so far.

    I had used 2.5-pro for a hefty chunk of development, and found it to work like a somewhat forgetful senior engineer who was new to my project. Very capable, would almost always choose a reasonable option, if not always the best one for the project, and not great at multi-tasking. Generally, made me comfortable not scrutinizing the code line-by-line, but still needed a bit of steering once projects got to a reasonable size.

    The next model is a clear step up in the multi-tasking capability at least, with me very rarely having to steer the implementation of a well-defined issue. In terms of code, I found MiMo-V.2.5-pro to be extremely conservative, implementing minimal solutions. The next model seems a little bit more ambitious, in positive ways, making good guesses about gaps/next steps. It also seems to be a fair bit better at design, at least for the little bit I've done, it was good at translating my concepts to practical elements on screen, and cleaned things up nicely as I made suggestions.

  4. ricardobeat

    For reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great.

    Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).

    https://deepswe.datacurve.ai/blog/deepswe-v1-1

  5. krm01

    This is pretty neat. What would be a good reason for the other Model providers to not do this?

  6. liuliu

    When you run benchmarks while training, isn't that the definition of contamination? Asking because I am not sure if this is normal in big labs now.

  7. fzysingularity

    Very cool to see the openness here, and likely more like this will come from smaller startups where they win users on transparency.

  8. ssn2000

    Total run cost is $1.2M until now, what resources are they using to train their model? Wish they shared more details on that and what the MFU metrics are.

More from this day

2026-09-16