I am increasingly hesitant to use non-native harnesses - model providers are now starting to train their agents for use within the harness. An eval like terminal bench can only capture so much data. I don't want to have to assess each harness every model release to make sure it's working as well as it can.
MiMo _needed_ it. MiMo 2.5 Pro really struggled when given more tools, causing it to lose other capabilities. e.g. in security benchmarks I was doing, most of the medium-to-large models got better or stayed roughly the same at finding security vulnerabilities when given more tools (e.g. treesitter, semgrep, a full bash/python environment, etc.) vs. when only given the ability to read the files in the repo. But, MiMo got notably worse. It seemed to get confused by all the options. MiMo 2.5 Pro just reading files is an excellent security bug finder, at the pareto frontier at the time (cheapest option to find as many bugs as Opus 4.8, which was current at the time), but adding tools cratered it down to small model territory.
I haven't tested the new version, and I still have a bunch of tokens on my token plan, so I might give 2.6 a go in some similar tests to see if it still has the analysis paralysis problem of 2.5.
johnmlussier · · focus · HN ↗
verdverm · · focus · HN ↗
<a href="https://mimo.mi.com/docs/en-US/news/latest/v2-6" rel="nofollow">https://mimo.mi.com/docs/en-US/news/latest/v2-6
SwellJoe · · focus · HN ↗
I haven't tested the new version, and I still have a bunch of tokens on my token plan, so I might give 2.6 a go in some similar tests to see if it still has the analysis paralysis problem of 2.5.
throwa356262 · · focus · HN ↗
The non-pro version is mostly useless
SwellJoe · · focus · HN ↗