I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve.
The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.
--
PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.
I've found that mimo v2.5 works for very basic things like a python script to do one thing, but it also is very 'dumb' compared to qwen 3.8-flash-next (I think the benchmark scores for terminal and coding specific benches back this up). And definitely not in the same class as like a GLM5.2 or 5.3. It's fast but makes basic mistakes that only get caught later.
The fact I can run Qwen 3.8 Flash Next locally, forever (on my DGX Spark-alike) is genuinely shocking to me. It’s crazy good for how small it is. Fast, too.
Yeah, I'm guessing you have a variant that fits in <128GB with 262k context? I have the unsloth Q8 GGUF of it here in a setup that with full context and ton of extra llama-server "--cache-ram" sits around 200GB RAM usage on a 256GB system, it's probably the best thing I've found for a 256GB class machine. Enough headroom for a rope/yarn extension to 524288 context if I need it.
RTX6000 Blackwell with 96GB is enough to run it with NV4, 256k context, KVcache, multimodal at 130t/s (SGLang). It's toasty, you're using up 94GB of those 96, but it works and the results are great
I made this 3D game in a day on the same setup with Qwen Code as agent: <a href="https://games.jonathanpage.com/" rel="nofollow">https://games.jonathanpage.com/
And I am not a web developer! It's an extraordinary model.
Yeah, this Brazilian dude who has been a contributor here on HN longer than your anonymous account is shilling for a Chinese model company. Makes sense.
They both are in the 50-100 tok/s range. The Mimo v2.5 Pro Ultraspeed beta could reach 1000 tok/s, hoping they can do something similar for the new model, it was amazing.
What was meant is probably opencode. It's free there, along with several other models. You have to use their harness to access it. It's alright.
mimo 2.5 has been a big underperformer since shortly after it's release imo. i cancelled my sub after the first month. purposefully using 2.5 right now is just handicapping yourself for no reason.
Same here. It’s the first AI provider I actually gave money to, since they offered the model for free with a Mimo code for the first month or so, and it was great.
These days, there are more intelligent models like DS4.1, but Mimo is very obedient, so I plan things with another model and give the implementation to Mimo.
Which models do you like for planning? IME I haven't seen much success with using cheap/open models for planning, so I use Claude for planning still.
For comparison, I am currently at 6.6B tokens, 95% of monthly quota on a 10$ command code plan, mostly using DeepSeek flash 4.1, or some of the free models for easier tasks.
That's an eternity when it comes to coding models.
In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.
That's because Opus 4.6 was the last good assistant model.
Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant.
Now it's *you* being the assistant, reviewer, etc.
Because they’ve essentially exhausted pre training scaling and are looking to post training to expand capabilities, which is really just optimization via reinforcement learning against specific tasks aka bench maxing.
In another one, Opus 4.6 level already solved 90% of my work-day tasks, so while better models have been instrumental into handling a higher % that does not mean that defaulting on cheaper models can't be good.
I run DS 4.1 flash daily, and then cross check with gpt-6-astra and I've nuked 90% of my AI monthly bill while having higher limits and better performance/intelligence than I did just at the beginning of this summer.
You use the low cost Mimo-V2.5 and not its big brother Mimo-V2.5 Pro? I also made good experience with Mimi-V2.5 when used in conjunction with prewalk mode. But then other models got so cheap and perform better, so I only use Mimo for background tasks.
The availability of Mimo over Openrouter got however, much worse recently.
joelwallis · · focus · HN ↗
The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.
-- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.
james2doyle · · focus · HN ↗
I always found that those Mimo models to be really good at tool calling and following instructions
walrus01 · · focus · HN ↗
[deleted] · · focus · HN ↗
[deleted]
girvo · · focus · HN ↗
walrus01 · · focus · HN ↗
girvo · · focus · HN ↗
It’s good enough that I’m considering a second spark, or selling this and buying an M5 Ultra with 256GB for it
gmerc · · focus · HN ↗
jonsoft · · focus · HN ↗
And I am not a web developer! It's an extraordinary model.
(Mouse and keyboard required)
jwpapi · · focus · HN ↗
eli · · focus · HN ↗
yeeeloit · · focus · HN ↗
[dead]
platinumrad · · focus · HN ↗
senordevnyc · · focus · HN ↗
esafak · · focus · HN ↗
ricardobeat · · focus · HN ↗
alwinaugustin · · focus · HN ↗
rahmatawaludin · · focus · HN ↗
farlight · · focus · HN ↗
<a href="https://opencode.ai/docs/zen/#pricing">https://opencode.ai/docs/zen/#pricing
flexagoon · · focus · HN ↗
electroglyph · · focus · HN ↗
[dead]
NuclearPM · · focus · HN ↗
electroglyph · · focus · HN ↗
NuclearPM · · focus · HN ↗
rapind · · focus · HN ↗
I’ll give Mimo a try.
trollbridge · · focus · HN ↗
UltraSpeed was absolutely awesome. I miss it.
DS 4.1 Flash is amazing. Well worth the extra cost.
miyuru · · focus · HN ↗
These days, there are more intelligent models like DS4.1, but Mimo is very obedient, so I plan things with another model and give the implementation to Mimo.
abustamam · · focus · HN ↗
baxtr · · focus · HN ↗
ignoramous · · focus · HN ↗
API may be expensive, but I do 900m tokens (95% cached, ~0.4% output) on Z.ai's $18/mo coding plan with GLM 5.3 Flash.
miroljub · · focus · HN ↗
For comparison, I am currently at 6.6B tokens, 95% of monthly quota on a 10$ command code plan, mostly using DeepSeek flash 4.1, or some of the free models for easier tasks.
epolanski · · focus · HN ↗
farlight · · focus · HN ↗
wangxili1997 · · focus · HN ↗
[dead]
ehsankia · · focus · HN ↗
That's an eternity when it comes to coding models.
In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.
pelagicAustral · · focus · HN ↗
epolanski · · focus · HN ↗
Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant.
Now it's *you* being the assistant, reviewer, etc.
pelagicAustral · · focus · HN ↗
transdev12 · · focus · HN ↗
ctolsen · · focus · HN ↗
epolanski · · focus · HN ↗
In another one, Opus 4.6 level already solved 90% of my work-day tasks, so while better models have been instrumental into handling a higher % that does not mean that defaulting on cheaper models can't be good.
I run DS 4.1 flash daily, and then cross check with gpt-6-astra and I've nuked 90% of my AI monthly bill while having higher limits and better performance/intelligence than I did just at the beginning of this summer.
sinuhe69 · · focus · HN ↗
The availability of Mimo over Openrouter got however, much worse recently.