I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.
The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).
If you’re releasing an open model going forward, please consider offering the community more of this transparency!
China will most probably win the AI race in the long run because of one major bottleneck the US has - energy. The electric energy and grid buildout in China has been massive since a long time and there is simply no way for the US to quickly catch up.
No matter how much cash you throw you can't just materialize a 100 nuclear reactors to power the data centers.
Looking at the frontend design examples; why do these models seem to love the "01 - UPPERCASE TEXT" motif. It's everywhere now (see https://try.cloudflare.com/, which has '01 · QUICK TUNNELS', but no "02" anywhere).
I said this years ago, LLMs are a commodity (or were becoming one at the time). They are dime a dozen. Even the frontier ones. OpenAI and Anthropic have no moat.
No moat and competition is good for consumers though.
Can someone recommend me a local harness to explore the multimodality of this kind of models? Text in text out is mostly generically solved, but that's about what I know
Mimo has been one of those models that I have been rooting for since the first I used it, the 2.5 pro which I have used quite a bit, was very concise, very aware of how much context needs to be read for which tasks and would always keep the context tight. Also surprisingly good at strategic thinking. I had published a comparison between it and Terra where Terra was found to be using much more avg context for similar tasks https://dirac.run/posts/gpt-5-6-vs-mimo-2-5-pro-context-bloa...
Interesting but not surprising trend across the board seems to be, the flash models seems to have caught up with the pro-sized models of H1'26. No surprise all labs are rushing to bigger models.
EDIT: Wow, took a detailed look at the benchmarks. Mimo 2.6 pro, the 1T model leads Kimi K3, a 2.8T param model in 14 out of 15 benchmarks (and the last one is near tie)!! Good to see they also kept the price the same, and landed in the greenest quardrant of the intelligence vs speed of AA.
Mimo has long been one of my preferred models. It works well at many code tasks, generally has a pleasant voice, and isn't prone to over-analzing and researching
Also when I was using it, I managed to do quite a bit on a few bucks worth of OpenRouter credits. Not sure how well it keeps up in the modern world against things like Luna, but I hope it remains competitive
Are folks working in biology / cybersecurity seeing more limits in what Opus (not Fable) is allowing? This has happened quite suddenly for me and I’m stuck in the middle of a project that would have otherwise called for use of Claude.
I’ll be trying these models out and may end up switching my subscriptions if this craziness continues
I tried them, but could only test Pro none and Flash none and low, the other ones (medium/high) used way too many tokens and all requests timed out. Not sure if they have a problem with their API, or this model is really token inefficient/basically unusable.
All these new models are such tease for us folks with 128GB of shared memory. Buying another unit now to expand to 256GB is a mortgage payment but it’s getting tempting…
Are people using MiMo models as their daily drivers in a company setting? If so, how? I know they're available over OpenCode and directly from Xiaomi, but those are not great options. OpenCode Go straight up doesn't give any guarantees about training on your data, and Xiaomi says that they won't but it's unclear.
With some models you can find hosting companies based in the EU or US, but then you don't know how they're quantizing the models, so you're not sure about the actual output quality.
How are people actually using this? Or are people just experimenting with side projects?
Overall this felt like the weakest of the three. Ultraspeed was fast as far as tokens per second goes, but it overthought quite a lot of things resulting it in having one of the longest build times. That overthinking didn't lead to better results either - note the statue with the cropped off arm. It's also the worst implementation of the dynamic lighting effect / displacement effect - the background especially has some significant distortion. Astra was the only one that seemed to understand that displacement should happen less the further something is in the distance.
is there a good metric of model degredation over time?
Im a bit lazy and only use the free models different companies host and the biggest difference i see is that some models (Gemini, OpenAI) get progressively stupid in long chats. You end up having to start a new session every oncr in a while. Or they get really hung up on a theme and cant shift to a new topic.
By contrast, Ive been impressed with Qwen. I have some chats on research and code architecture that have stretched for weeks without any noteable change in quality (though occassionally it seems to "rush" to an answer)
Im just looking at all the listed benchmarks and im unsue which i should be looking at
Please, think about the american companies!
If we don't protect them, when they go broke (because they will, it will be the perfect triple dip) and we bail them with american taxes we will have paid twice for all the work they had already stealed (not my words, this is a microsoft anti-ai exec...)
Twice, trice or quadrix(tm) are just approximations, we have already paid with:
- more expensive electronics
- less work
- all the retirement money put into gpus
- all the "fair use" of all the books, all the images,
- and then taxes to bail them?
Any idea on how they get the pricing so efficient? Their artificialanalysis graph has them on par with GLM 5.3 but at less than 10% of the price despite being larger than GLM 5.3.
Anyone knows what they used to create the videos ? Is the model driving a program like Davinci Resolve / After Effects ? or is the model writing code to then generate these videos via some library.
Granted this is an awesome release and I loved watching the livestreamed RL dashboard, I found this message on the dashboard (https://mimo.xiaomi.com/rl/) quite funny:
> we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs.
And this in the model card (emphasis mine):
> Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.
Environment hardening during the RL runs. Uh oh, did someone start making a few too many paperclips?
I really liked MiMo 2.5, it was really affordable and actually had vision, unlike DeepSeek. (DeepSeek has only recently added it)
Just tried 2.6 flash on a really niche topic I specialise in and it has done a really good job. They’ve definitely polluted their training data with claudeslop, but looking past the slop there is a decent model.
For simple edits it produces huge reasoning traces. Kind of disappointing. Super long repetitive reasoning always sits wrong with me. It feels like a way for the labs to brute-force higher benchmark scores but not actually like a smarter model. Disappointing, Mimo2.5 was such a nice model.
Funny that all but one video has audio, the house 3D model one, where you can hear (what I assume are) Xiaomi's engineers talking about who knows what.
This is a big week. Probably getting next OpenAI and Anthro models, Grok 4.7, Mimo, etc. These open source model releases are why I can't take the "slow down" crowd seriously. I pitted older Mimo, qwen, step, gpt-oss, and other models against each other playing games like Werewolf and Sketch.io-like games where I let them talk shit while they played against each other. Mimo was by far pareto frontier of game-playing for the models that were <$0.15/m input tokens on OpenRouter. Qwen was pareto frontier in the shit talking game though. Qwen's hilarious. https://www.tiktok.com/@clankerfights/video/7642862917582425...
Is there a way to test this in a browser interface? I'm not going to run any executables from any of the frontier labs on my machine, be that Xiaomi or Google or Anthropic.
The moat for OAI and anthropic seems to be very quickly shrinking. Chinese labs are now using RSI-like approaches and even without resorting to heavy distillation they're catching up in a couple of months vs. what would have been 6-12 months a year prior.
And as these models get better the pace of training is quickly speeding up too.
This doesn't bode particularly well for anthropic/OAI after they go public.
Leaning into what it cost to train is hilarious and an obvious shot at US frontier labs spending tens to hundreds of millions or more to train their models.
MiMo-V2.6-Flash-310B-A15B roughly GPT-5.6 Luna / Claude 4.9 according to benchmarks
MiMo-V2.6-Pro-1.02T-A42B roughly GPT-5.6 Sol / Opus 5 according to benchmarks
Experiences so far with the non-flash model: a bit of an overthinker/overplanner. Makes bad/messy architecture choices with an open-ended prompt. But actually very competent at intricate code level things, and pretty good at finding corner cases, including in code that other more advanced models had looked at and not noticed problems with.
I deliberately left things wide open and ambiguous to see what it would do. I think with better upfront planning this would be excellent value.
ah, would you look at that. I was wondering why mimo 2.5 became "dumber" the last weeks. I was speculating they are probably about to release a new version of the model. because the model really acted out a lot. especially the last two weeks. dont know, was just a feeling, highly speculative.
> This marks a key step in our exploration of the RSI path: scaling RL compute on verifiable, complex tasks, so the model can continuously expand its capability frontier through exploration and feedback.
Gotta love a capable open model. BUT, how can they just casually throw in that they're actively exploring RSI as if it's just another technique? Is this not alarming at all?
The Xiaomi always somehow lacks the quality of competitors. mobile phones are lower quality or weird. the login reset mechanism sucks. If I can't even login, how good this system is supposed to be.
They are excellent in marketing, I guess that is something.
In the chart they use "Pareto Line", which I think is wrong. Pareto is 20% effort leading to 80% results. Which could be interpreted as models costing 20% having 80% of peak intelligence, but that’s not what it looks like to me.
It looks like the "Frontier Line" to me, which is also often misinterpreted. frontier does not mean the best models. It means all models that are not strictly dominated, meaning in most cases: Not same price or cheaper and more intelligent.
I personally would like the word frontier to be used with more criterias: Open Weights, per use-case, etc etc. This would make model selection easier, but I understand it’s not an easy thing to do.
MiMo v2.6
(mimo.xiaomi.com)924 points by volf_ 14 hours ago | 413 comments
Comments
The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).
If you’re releasing an open model going forward, please consider offering the community more of this transparency!
No matter how much cash you throw you can't just materialize a 100 nuclear reactors to power the data centers.
Pro [2]:, 1.02T total / 42B activated parameters
[1]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
[2]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
Maybe Terminal Bench 4.0 and ExploitGym are reasonable.
Terminal Bench 4.0
ExploitGym DeepSWE v1.1https://aibenchy.com/compare/x-ai-grok-4-7-medium/xiaomi-mim...
Pelicans for Pro: https://tools.simonwillison.net/markdown-svg-renderer?url=ht...
No moat and competition is good for consumers though.
Some features of the release I like:
- Demonstration of diverse tasks, such as using a DAW
- Graphs from various benchmarks and price ranges
- Real world use of the model in scientific environments
Averages ~25-35tok/s which isn't bad for a first attempt.
Interesting but not surprising trend across the board seems to be, the flash models seems to have caught up with the pro-sized models of H1'26. No surprise all labs are rushing to bigger models.
EDIT: Wow, took a detailed look at the benchmarks. Mimo 2.6 pro, the 1T model leads Kimi K3, a 2.8T param model in 14 out of 15 benchmarks (and the last one is near tie)!! Good to see they also kept the price the same, and landed in the greenest quardrant of the intelligence vs speed of AA.
Also when I was using it, I managed to do quite a bit on a few bucks worth of OpenRouter credits. Not sure how well it keeps up in the modern world against things like Luna, but I hope it remains competitive
It's because offpeak electricity is cheaper?
Funnily it's perfect if you are in the Pacific Time Zone because you can use it daytime 9am to 5pm
I’ll be trying these models out and may end up switching my subscriptions if this craziness continues
With some models you can find hosting companies based in the EU or US, but then you don't know how they're quantizing the models, so you're not sure about the actual output quality.
How are people actually using this? Or are people just experimenting with side projects?
Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...
MiMo 2.6 Pro Ultraspeed (36min): https://html.non.io/annui-mimo/
Grok 4.7 (25min): https://html.non.io/Annui-grok/
Astra (19min): https://html.non.io/annui/
Overall this felt like the weakest of the three. Ultraspeed was fast as far as tokens per second goes, but it overthought quite a lot of things resulting it in having one of the longest build times. That overthinking didn't lead to better results either - note the statue with the cropped off arm. It's also the worst implementation of the dynamic lighting effect / displacement effect - the background especially has some significant distortion. Astra was the only one that seemed to understand that displacement should happen less the further something is in the distance.
Here's a vid of all 3 side by side with the source design: https://non.io/video/annui-comparison.mp4
Staggeringly low for a frontier model.
Im a bit lazy and only use the free models different companies host and the biggest difference i see is that some models (Gemini, OpenAI) get progressively stupid in long chats. You end up having to start a new session every oncr in a while. Or they get really hung up on a theme and cant shift to a new topic.
By contrast, Ive been impressed with Qwen. I have some chats on research and code architecture that have stretched for weeks without any noteable change in quality (though occassionally it seems to "rush" to an answer)
Im just looking at all the listed benchmarks and im unsue which i should be looking at
Twice, trice or quadrix(tm) are just approximations, we have already paid with:
KillSwitch-Bench 1.0
1 - https://bench.killswitch-lang.org/> we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs.
And this in the model card (emphasis mine):
> Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.
Environment hardening during the RL runs. Uh oh, did someone start making a few too many paperclips?
Just tried 2.6 flash on a really niche topic I specialise in and it has done a really good job. They’ve definitely polluted their training data with claudeslop, but looking past the slop there is a decent model.
And as these models get better the pace of training is quickly speeding up too.
This doesn't bode particularly well for anthropic/OAI after they go public.
it is reasonable to question why America doesn't have such environment.
https://artificialanalysis.ai/models/mimo-v2-6-pro#intellige...
That's pretty fast; I think I'll try it: https://openrouter.ai/xiaomi/mimo-v2.6-flash
One concern I have is that they allegedly do not discount cached tokens: https://www.reddit.com/r/opencodeCLI/comments/1t37dz3/xiaomi...
Can anyone comment?
so weird to acknowledge someone being on the front edge, but not name it
MiMo-V2.6-Flash-310B-A15B roughly GPT-5.6 Luna / Claude 4.9 according to benchmarks MiMo-V2.6-Pro-1.02T-A42B roughly GPT-5.6 Sol / Opus 5 according to benchmarks
Perhaps with IQ2 flash will run on 128G M5?
I deliberately left things wide open and ambiguous to see what it would do. I think with better upfront planning this would be excellent value.
but now I got my "proof".
Gotta love a capable open model. BUT, how can they just casually throw in that they're actively exploring RSI as if it's just another technique? Is this not alarming at all?
They are excellent in marketing, I guess that is something.
It looks like the "Frontier Line" to me, which is also often misinterpreted. frontier does not mean the best models. It means all models that are not strictly dominated, meaning in most cases: Not same price or cheaper and more intelligent.
I personally would like the word frontier to be used with more criterias: Open Weights, per use-case, etc etc. This would make model selection easier, but I understand it’s not an easy thing to do.