‹ BackHN Continuity

Thread

Xiaomi Mimo 2.6 live post-training dashboard

562 points · 155 comments · krackers

  1. ricardobeat · · focus · HN ↗
    For reference, Mimo-v2.5-Pro scored 19% on DeepSWE 1.1. This is looking great.

    Fable scores 70%, Kimi K3 69%, Astra 74% (all on max effort).

    <a href="https:&#x2F;&#x2F;deepswe.datacurve.ai&#x2F;blog&#x2F;deepswe-v1-1">https:&#x2F;&#x2F;deepswe.datacurve.ai&#x2F;blog&#x2F;deepswe-v1-1

    1. Cookingboy · · focus · HN ↗
      2.6-pro just reached 63.7% by step 10, it&#x27;s on step 11 right now.

      Even flash reached 60.7% by step 12, and it&#x27;s on step 16 now.

      This is so exciting lmao.

      1. arcanemachiner · · focus · HN ↗
        DeepSWE is saturated now IMO, and is basically worthless. Lots of new models get around 74%. Shame too, because it was a pretty decent benchmark for a few months there.
        1. brookst · · focus · HN ↗
          It is saturated, but that doesn’t mean worthless. Seeing 72% is low-signal, but 30% is still meaningful.
    2. markasoftware · · focus · HN ↗
      gemini 3.8 flash is also 74% and google just started letting all their engineers use claude...go figure
      1. ehsankia · · focus · HN ↗
        &gt; and google just started letting all their engineers use claude

        That&#x27;s misleading.

        1. Having different models available is useful for A&#x2F;B testing and helping improve Gemini itself.

        2. They have an enterprise offering for Antigravity (their agentic coding platform), and they need to test that it works well with non-Gemini models too.

    3. buffalobuffalo · · focus · HN ↗
      Also worth taking a look at is the mimo harness. It&#x27;s a fork of opencode with some new modes added for long horizon tasks. One of the better open harnesses out there at the moment.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.