‹ BackHN Continuity

Thread

Xiaomi Mimo 2.6 live post-training dashboard

562 points · 155 comments · krackers

  1. joelwallis · · focus · HN ↗
    I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve.

    The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.

    -- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.

    1. ehsankia · · focus · HN ↗
      > late last year/early this year

      That's an eternity when it comes to coding models.

      In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.

      1. pelagicAustral · · focus · HN ↗
        tbf, I the happiest I've been working with claude is late last year/early this year (before March)...
        1. epolanski · · focus · HN ↗
          That's because Opus 4.6 was the last good assistant model.

          Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant.

          Now it's *you* being the assistant, reviewer, etc.

          1. pelagicAustral · · focus · HN ↗
            Is that right? Why the shit would that happen?
            1. transdev12 · · focus · HN ↗
              Because they’ve essentially exhausted pre training scaling and are looking to post training to expand capabilities, which is really just optimization via reinforcement learning against specific tasks aka bench maxing.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.