‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. dom96 · · focus · HN ↗
    Very capable model. I just ran it on my own LLM benchmark suite[1] and it matches Muse Spark 1.3 in pass rate but is significantly cheaper.

    KillSwitch-Bench 1.0

      Claude Opus 5           66.9
      GPT-6 Astra             57.9
      Claude Fable 5.1        46.7
      MiMo-V2.6-Pro           38.8
      Muse Spark 1.3          36.5
    
    1 - <a href="https:&#x2F;&#x2F;bench.killswitch-lang.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;bench.killswitch-lang.org&#x2F;
    1. bel8 · · focus · HN ↗
      I appreciate that this benchmark is different but it is in no way how most people use LLMs or promote equal grounds when benchmarking:

      - capped per-task budget and time limit

      - No internet access

      - different harnesses mixed

      1. yt1998 · · focus · HN ↗

        [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.