‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. dom96 · · focus · HN ↗
    Very capable model. I just ran it on my own LLM benchmark suite[1] and it matches Muse Spark 1.3 in pass rate but is significantly cheaper.

    KillSwitch-Bench 1.0

      Claude Opus 5           66.9
      GPT-6 Astra             57.9
      Claude Fable 5.1        46.7
      MiMo-V2.6-Pro           38.8
      Muse Spark 1.3          36.5
    
    1 - <a href="https:&#x2F;&#x2F;bench.killswitch-lang.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;bench.killswitch-lang.org&#x2F;
    1. bel8 · · focus · HN ↗
      I appreciate that this benchmark is different but it is in no way how most people use LLMs or promote equal grounds when benchmarking:

      - capped per-task budget and time limit

      - No internet access

      - different harnesses mixed

      1. dom96 · · focus · HN ↗
        How do you think other benchmarks work? Every single one is going to have a budget and time limit. There has to be a cap on those, you can&#x27;t just let it spin forever and use unlimited funds.

        I would argue that this benchmark is uniquely suited to how most people use LLMs because it actually tests common harnesses and it is a true coding benchmark for a language that is unseen, thus testing the LLMs actual ability to understand nuance and learn.

        Internet access is restricted to ensure that over time models cannot just look up the source code of KillSwitch, which would allow them to cheat.

        1. bel8 · · focus · HN ↗
          I didn&#x27;t expect there to be a time limit and I do know benchmarks that don&#x27;t have a time limit. They just penalize slower models which I think is fair.

          In my experience models just don&#x27;t take forever to mark tasks as done.

          For my coding usage I don&#x27;t set time limit so benchmarks that do so are providing me less interesting use cases.

          With that said, I can understand having budget limits for expensive LLMs for those of us that don&#x27;t have infinite VC money.

          As for no internet access, the issue is that for most problems we do want Llms to be able to search docs, API specs, GitHub issues, etc. So not allowing that usually just favours larger LLMs that were able to memorize more data, not necessarily smarter ones when both have internet access.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.