‹ BackHN Continuity

Thread

Xiaomi Mimo 2.6 live post-training dashboard

562 points · 155 comments · krackers

  1. joelwallis · · focus · HN ↗
    I been using MiMo-V2.5 to do most of my work as software engineer, on a variety of projects I'm working on, and I been VERY happy with ROI. The model is very powerful! Not perfect – I've run in hallucination loops once or twice, but nothing a stop-then-continue wouldn't solve.

    The cost is unbelievably low, and the quality of intelligence I get is equivalent to when I was working mostly with Anthropic models (late last year/early this year). I'm fully invested in MiMo and I'm very happy with it.

    -- PS: I also check almost daily to see if other models are capable of doing such great work. And they do – DS4F is powerful and DS41 is impressive, GLM 5.3 Flash gets a job done well, etc. – but when I add cost of M-token in the ROI math, Jeez! MiMo is an order of magnitude better.

    1. ehsankia · · focus · HN ↗
      > late last year/early this year

      That's an eternity when it comes to coding models.

      In my personal experience, we've had almost a step change every ~3 months this year, at least for bigger one-shot tasks. For example looking at Gemini Flash 3.0 vs 3.5 vs 3.8, it went 5% -> 30% -> 75% on DeepSWE, all since the start of the year.

      1. pelagicAustral · · focus · HN ↗
        tbf, I the happiest I've been working with claude is late last year/early this year (before March)...
        1. epolanski · · focus · HN ↗
          That's because Opus 4.6 was the last good assistant model.

          Everything after it might be more "intelligent" but is super tuned around end-to-end task (and related benchmarks), not to act as an assistant.

          Now it's *you* being the assistant, reviewer, etc.

          1. pelagicAustral · · focus · HN ↗
            Is that right? Why the shit would that happen?
            1. ctolsen · · focus · HN ↗
              Their ambition isn't your work being amplified by their model, they want you running fifty autonomous long-running agents.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.