MiMo v2.6

(mimo.xiaomi.com)

Comments

rao-v 14 hours ago
I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.

The realtime dashboard they shared during training (https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).

If you’re releasing an open model going forward, please consider offering the community more of this transparency!

margorczynski 11 hours ago
China will most probably win the AI race in the long run because of one major bottleneck the US has - energy. The electric energy and grid buildout in China has been massive since a long time and there is simply no way for the US to quickly catch up.

No matter how much cash you throw you can't just materialize a 100 nuclear reactors to power the data centers.

stymaar 14 hours ago
Flash[1]: 309B total / 15B activated parameters

Pro [2]:, 1.02T total / 42B activated parameters

[1]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL

[2]: https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL

user43928 14 hours ago
I don't trust any of the benchmarks where Opus 5 surpasses Astra or Fable 5.1.

Maybe Terminal Bench 4.0 and ExploitGym are reasonable.

Terminal Bench 4.0

  GPT 6 Astra             59.6
  Claude Fable 5.1        55.1
  Claude Opus 5           49.0
  MiMo-V2.6-Pro           34.9
  MiMo-V2.6-Flash         28.8
  DeepSeek V4.1 Flash     26.8
  MiMo-V2.5-Pro            1.5
ExploitGym

  GPT 6 Astra             42.4
  Claude Fable 5.1        30.4
  Claude Opus 5           22.1
  MiMo-V2.6-Pro           17.8
  MiMo-V2.6-Flash          6.0
  MiMo-V2.5-Pro            0.1
DeepSWE v1.1

  DeepSeek V4.1 Flash     74.2
  Claude Opus 5           74.0
  GPT 6 Astra             74.0
  MiMo-V2.6-Pro           71.9
  Claude Fable 5          70.0
  MiMo-V2.6-Flash         67.9
  MiMo-V2.5-Pro           19.0
XCSme 3 hours ago
It seems to be better than Grok 4.7 and a lot cheaper in my tests:

https://aibenchy.com/compare/x-ai-grok-4-7-medium/xiaomi-mim...

nemothekid 14 hours ago
Looking at the frontend design examples; why do these models seem to love the "01 - UPPERCASE TEXT" motif. It's everywhere now (see https://try.cloudflare.com/, which has '01 · QUICK TUNNELS', but no "02" anywhere).
toephu2 12 hours ago
I said this years ago, LLMs are a commodity (or were becoming one at the time). They are dime a dozen. Even the frontier ones. OpenAI and Anthropic have no moat.

No moat and competition is good for consumers though.

vatsachak 14 hours ago
Wow, the chinese labs are getting good at advertising model releases. The moat is thin.

Some features of the release I like:

- Demonstration of diverse tasks, such as using a DAW

- Graphs from various benchmarks and price ranges

- Real world use of the model in scientific environments

SSLy 26 minutes ago
Can someone recommend me a local harness to explore the multimodality of this kind of models? Text in text out is mostly generically solved, but that's about what I know
volf_ 13 hours ago
I've got a working recipe to run this model on Dual DGX Spark: https://github.com/volfco/spark-vllm-docker/blob/main/recipe...

Averages ~25-35tok/s which isn't bad for a first attempt.

GodelNumbering 12 hours ago
Mimo has been one of those models that I have been rooting for since the first I used it, the 2.5 pro which I have used quite a bit, was very concise, very aware of how much context needs to be read for which tasks and would always keep the context tight. Also surprisingly good at strategic thinking. I had published a comparison between it and Terra where Terra was found to be using much more avg context for similar tasks https://dirac.run/posts/gpt-5-6-vs-mimo-2-5-pro-context-bloa...

Interesting but not surprising trend across the board seems to be, the flash models seems to have caught up with the pro-sized models of H1'26. No surprise all labs are rushing to bigger models.

EDIT: Wow, took a detailed look at the benchmarks. Mimo 2.6 pro, the 1T model leads Kimi K3, a 2.8T param model in 14 out of 15 benchmarks (and the last one is near tie)!! Good to see they also kept the price the same, and landed in the greenest quardrant of the intelligence vs speed of AA.

paradox460 10 hours ago
Mimo has long been one of my preferred models. It works well at many code tasks, generally has a pleasant voice, and isn't prone to over-analzing and researching

Also when I was using it, I managed to do quite a bit on a few bucks worth of OpenRouter credits. Not sure how well it keeps up in the modern world against things like Luna, but I hope it remains competitive

thrownawaysz 13 hours ago
>Night 0.8x Usage, 00:00-08:00 -UTC+8

It's because offpeak electricity is cheaper?

Funnily it's perfect if you are in the Pacific Time Zone because you can use it daytime 9am to 5pm

bonsai_spool 12 hours ago
Are folks working in biology / cybersecurity seeing more limits in what Opus (not Fable) is allowing? This has happened quite suddenly for me and I’m stuck in the middle of a project that would have otherwise called for use of Claude.

I’ll be trying these models out and may end up switching my subscriptions if this craziness continues

XCSme 11 hours ago
I tried them, but could only test Pro none and Flash none and low, the other ones (medium/high) used way too many tokens and all requests timed out. Not sure if they have a problem with their API, or this model is really token inefficient/basically unusable.
syntaxing 14 hours ago
All these new models are such tease for us folks with 128GB of shared memory. Buying another unit now to expand to 256GB is a mortgage payment but it’s getting tempting…
wkcheng 12 hours ago
Are people using MiMo models as their daily drivers in a company setting? If so, how? I know they're available over OpenCode and directly from Xiaomi, but those are not great options. OpenCode Go straight up doesn't give any guarantees about training on your data, and Xiaomi says that they won't but it's unclear.

With some models you can find hosting companies based in the EU or US, but then you don't know how they're quantizing the models, so you're not sure about the actual output quality.

How are people actually using this? Or are people just experimenting with side projects?

jjcm 12 hours ago
Here's an image->html test for it using 2.6 Pro Ultraspeed, along with comparisons for grok 4.7 and Astra.

Design: https://image.non.io/78795662-8bfc-4e14-8d72-3738392aa6b3.we...

MiMo 2.6 Pro Ultraspeed (36min): https://html.non.io/annui-mimo/

Grok 4.7 (25min): https://html.non.io/Annui-grok/

Astra (19min): https://html.non.io/annui/

Overall this felt like the weakest of the three. Ultraspeed was fast as far as tokens per second goes, but it overthought quite a lot of things resulting it in having one of the longest build times. That overthinking didn't lead to better results either - note the statue with the cropped off arm. It's also the worst implementation of the dynamic lighting effect / displacement effect - the background especially has some significant distortion. Astra was the only one that seemed to understand that displacement should happen less the further something is in the distance.

Here's a vid of all 3 side by side with the source design: https://non.io/video/annui-comparison.mp4

Alien1Being 3 hours ago
Training cost $ 3.47 million....

Staggeringly low for a frontier model.

geokon 5 hours ago
is there a good metric of model degredation over time?

Im a bit lazy and only use the free models different companies host and the biggest difference i see is that some models (Gemini, OpenAI) get progressively stupid in long chats. You end up having to start a new session every oncr in a while. Or they get really hung up on a theme and cant shift to a new topic.

By contrast, Ive been impressed with Qwen. I have some chats on research and code architecture that have stretched for weeks without any noteable change in quality (though occassionally it seems to "rush" to an answer)

Im just looking at all the listed benchmarks and im unsue which i should be looking at

xlayn 10 hours ago
Please, think about the american companies! If we don't protect them, when they go broke (because they will, it will be the perfect triple dip) and we bail them with american taxes we will have paid twice for all the work they had already stealed (not my words, this is a microsoft anti-ai exec...)

Twice, trice or quadrix(tm) are just approximations, we have already paid with:

   - more expensive electronics
   - less work
   - all the retirement money put into gpus
   - all the "fair use" of all the books, all the images, 
   - and then taxes to bail them?
Havoc 11 hours ago
3.5 mil cost for a opus level model? Even if excluding salaries that seems very cheap
wmedrano 10 hours ago
Any idea on how they get the pricing so efficient? Their artificialanalysis graph has them on par with GLM 5.3 but at less than 10% of the price despite being larger than GLM 5.3.
lwansbrough 14 hours ago
Anyone else more excited about Chinese models than American models these days? Big thing for me is affordability.
pulkitsh1234 12 hours ago
Anyone knows what they used to create the videos ? Is the model driving a program like Davinci Resolve / After Effects ? or is the model writing code to then generate these videos via some library.
dom96 12 hours ago
Very capable model. I just ran it on my own LLM benchmark suite[1] and it matches Muse Spark 1.3 in pass rate but is significantly cheaper.

KillSwitch-Bench 1.0

  Claude Opus 5           66.9
  GPT-6 Astra             57.9
  Claude Fable 5.1        46.7
  MiMo-V2.6-Pro           38.8
  Muse Spark 1.3          36.5
1 - https://bench.killswitch-lang.org/
informal007 10 hours ago
I'm thinking if this is a more fair comparison among other models, like ChatGPT, Claude and Deepseek
wren6991 7 hours ago
Granted this is an awesome release and I loved watching the livestreamed RL dashboard, I found this message on the dashboard (https://mimo.xiaomi.com/rl/) quite funny:

> we also removed the cyber dataset from the upcoming pro run, since we observed some bad patterns in the rollout logs.

And this in the model card (emphasis mine):

> Aligned RL: Cold start from self-correction — the model reflects on and rewrites its own misaligned turns into grounded next steps. Throughout RL, environment hardening, adversarial screening, and verifier cross-checks keep the loop honest against reward hacking.

Environment hardening during the RL runs. Uh oh, did someone start making a few too many paperclips?

ddxv 14 hours ago
This looks great in terms of cost and capabilities, truly pushing the frontier forward in terms of open weight light weight models.
puszczyk 3 hours ago
Very informative page vs the recent grok release
MisterMunchkin 14 hours ago
I really liked MiMo 2.5, it was really affordable and actually had vision, unlike DeepSeek. (DeepSeek has only recently added it)

Just tried 2.6 flash on a really niche topic I specialise in and it has done a really good job. They’ve definitely polluted their training data with claudeslop, but looking past the slop there is a decent model.

est 9 hours ago
This model & pricing fits LeiJun's visio for xiaomi: The costco wholesale of tech companies.
Imanari 2 hours ago
For simple edits it produces huge reasoning traces. Kind of disappointing. Super long repetitive reasoning always sits wrong with me. It feels like a way for the labs to brute-force higher benchmark scores but not actually like a smarter model. Disappointing, Mimo2.5 was such a nice model.
eriquesito 13 hours ago
Funny that all but one video has audio, the house 3D model one, where you can hear (what I assume are) Xiaomi's engineers talking about who knows what.
DanMcInerney 14 hours ago
This is a big week. Probably getting next OpenAI and Anthro models, Grok 4.7, Mimo, etc. These open source model releases are why I can't take the "slow down" crowd seriously. I pitted older Mimo, qwen, step, gpt-oss, and other models against each other playing games like Werewolf and Sketch.io-like games where I let them talk shit while they played against each other. Mimo was by far pareto frontier of game-playing for the models that were <$0.15/m input tokens on OpenRouter. Qwen was pareto frontier in the shit talking game though. Qwen's hilarious. https://www.tiktok.com/@clankerfights/video/7642862917582425...
stemlord 11 hours ago
Stupid question: in the benchmark diagrams I'm assuming the values are percentiles, so what does 100% represent?
algoth1 14 hours ago
Finally a lab that doesn't cheat on the charts
heyjstn 8 hours ago
How could the flash model beat the pro on the Cyber benchmark?
drob518 12 hours ago
Conspicuous that there’s no reference to GLM 5.3/Flash in the reported benchmarks. Just Deepseek and Kimi.
jonathanstrange 2 hours ago
Is there a way to test this in a browser interface? I'm not going to run any executables from any of the frontier labs on my machine, be that Xiaomi or Google or Anthropic.
coss 11 hours ago
Anyone know what game engine its using to make the 3d game?
nivance 8 hours ago
it sounds greate. I still have 50% of my quota this month, so I'll continue renewing to give it a try.
alfalfasprout 13 hours ago
The moat for OAI and anthropic seems to be very quickly shrinking. Chinese labs are now using RSI-like approaches and even without resorting to heavy distillation they're catching up in a couple of months vs. what would have been 6-12 months a year prior.

And as these models get better the pace of training is quickly speeding up too.

This doesn't bode particularly well for anthropic/OAI after they go public.

system2 9 hours ago
As I read this I am getting API overloaded errors from Claude. I will switch to something else soon. I hate Claude.
tw1984 3 hours ago
it is pretty amazing to see a young & talented woman is the tech lead of this frontier model.

it is reasonable to question why America doesn't have such environment.

bertili 14 hours ago
They mixed up DeepSeek 4.1 Flash with something else on this page, possibly DeepSeek 4.1 Flash means Gemini 3.8 Flash.
esafak 12 hours ago
It tops the intelligence vs cost Pareto frontier and, uniquely for a Chinese model, does well in response time too.

https://artificialanalysis.ai/models/mimo-v2-6-pro#intellige...

That's pretty fast; I think I'll try it: https://openrouter.ai/xiaomi/mimo-v2.6-flash

One concern I have is that they allegedly do not discount cached tokens: https://www.reddit.com/r/opencodeCLI/comments/1t37dz3/xiaomi...

Can anyone comment?

gigatexal 13 hours ago
Leaning into what it cost to train is hilarious and an obvious shot at US frontier labs spending tens to hundreds of millions or more to train their models.
NooneAtAll3 14 hours ago
does anyone know what unnamed model is on paretto frontier picture right between MiMo 2.5 and 2.6?

so weird to acknowledge someone being on the front edge, but not name it

varispeed 13 hours ago
These benchmark are useless as they don't say whether they were done before or after Fable and Astra got nerfed.
spwa4 14 hours ago
As for the stats that everyone wants:

MiMo-V2.6-Flash-310B-A15B roughly GPT-5.6 Luna / Claude 4.9 according to benchmarks MiMo-V2.6-Pro-1.02T-A42B roughly GPT-5.6 Sol / Opus 5 according to benchmarks

Perhaps with IQ2 flash will run on 128G M5?

cmrdporcupine 11 hours ago
Experiences so far with the non-flash model: a bit of an overthinker/overplanner. Makes bad/messy architecture choices with an open-ended prompt. But actually very competent at intricate code level things, and pretty good at finding corner cases, including in code that other more advanced models had looked at and not noticed problems with.

I deliberately left things wide open and ambiguous to see what it would do. I think with better upfront planning this would be excellent value.

omani 14 hours ago
ah, would you look at that. I was wondering why mimo 2.5 became "dumber" the last weeks. I was speculating they are probably about to release a new version of the model. because the model really acted out a lot. especially the last two weeks. dont know, was just a feeling, highly speculative.

but now I got my "proof".

Zaraif13 8 hours ago
> This marks a key step in our exploration of the RSI path: scaling RL compute on verifiable, complex tasks, so the model can continuously expand its capability frontier through exploration and feedback.

Gotta love a capable open model. BUT, how can they just casually throw in that they're actively exploring RSI as if it's just another technique? Is this not alarming at all?

paperboy10000 8 hours ago
The Xiaomi always somehow lacks the quality of competitors. mobile phones are lower quality or weird. the login reset mechanism sucks. If I can't even login, how good this system is supposed to be.

They are excellent in marketing, I guess that is something.

jwpapi 14 hours ago
In the chart they use "Pareto Line", which I think is wrong. Pareto is 20% effort leading to 80% results. Which could be interpreted as models costing 20% having 80% of peak intelligence, but that’s not what it looks like to me.

It looks like the "Frontier Line" to me, which is also often misinterpreted. frontier does not mean the best models. It means all models that are not strictly dominated, meaning in most cases: Not same price or cheaper and more intelligent.

I personally would like the word frontier to be used with more criterias: Open Weights, per use-case, etc etc. This would make model selection easier, but I understand it’s not an easy thing to do.