‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. user43928 · · focus · HN ↗
    I don't trust any of the benchmarks where Opus 5 surpasses Astra or Fable 5.1.

    Maybe Terminal Bench 4.0 and ExploitGym are reasonable.

    Terminal Bench 4.0

      GPT 6 Astra             59.6
      Claude Fable 5.1        55.1
      Claude Opus 5           49.0
      MiMo-V2.6-Pro           34.9
      MiMo-V2.6-Flash         28.8
      DeepSeek V4.1 Flash     26.8
      MiMo-V2.5-Pro            1.5
    
    ExploitGym

      GPT 6 Astra             42.4
      Claude Fable 5.1        30.4
      Claude Opus 5           22.1
      MiMo-V2.6-Pro           17.8
      MiMo-V2.6-Flash          6.0
      MiMo-V2.5-Pro            0.1
    
    DeepSWE v1.1

      DeepSeek V4.1 Flash     74.2
      Claude Opus 5           74.0
      GPT 6 Astra             74.0
      MiMo-V2.6-Pro           71.9
      Claude Fable 5          70.0
      MiMo-V2.6-Flash         67.9
      MiMo-V2.5-Pro           19.0
    1. mokre · · focus · HN ↗
      Maybe you should not trust any of the benchmarks!
    2. varispeed · · focus · HN ↗
      They match my experience. Astra and Fable I rate below Sonnet. They are incredibly poor. They were excellent for a couple of days after release and then plummeted.

      Maybe I am being routed to more quantised versions or less capable models with system prompt to fake Astra or Fable.

      1. [deleted] · · focus · HN ↗

        [deleted]

    3. dom96 · · focus · HN ↗
      Why not? In my own benchmark Opus 5 does in fact come out on top[1]

      1 - <a href="https:&#x2F;&#x2F;bench.killswitch-lang.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;bench.killswitch-lang.org&#x2F;

      1. user43928 · · focus · HN ↗
        Good question, maybe I am underestimating it based on its absolutely horrible writing style.
        1. cosmojg · · focus · HN ↗
          Yeah, for better or worse, writing style is practically uncorrelated with agentic performance, which is all the rage right now and the thing that most popular benchmarks currently prioritize.
    4. novaleaf · · focus · HN ↗
      can you recommend any benchmark websites that show up-to-date details like this?

      TerminaBench, DeepSwe sites are out of date.

      1. UnfitFootprint · · focus · HN ↗
        Yeah it’s a shame a lot of these benchmarks are behind. My favourite was ‘SlopCodeBench’ [1] as I’m most interested in ai reinforcing its own bad decisions, but it’s not even up to current gen oai

        1: <a href="https:&#x2F;&#x2F;www.scbench.ai&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.scbench.ai&#x2F;

      2. 1899-12-30 · · focus · HN ↗
        <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;evaluations&#x2F;terminalbench-4-0" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;evaluations&#x2F;terminalbench-4-0
    5. 3abiton · · focus · HN ↗
      We&#x27;re past the one model fits them all kind of LLM. Most of the recent release actually regress on world knowledge for example, but optimize for something different: tool usage, thinking process, and agentic approach. And yes, in my own usage, some usecases Opus beats Fable.
    6. HighGoldstein · · focus · HN ↗
      I think we are still far from nailing down good LLM benchmarks, because the more general-purpose your software the harder the question of what makes it good becomes. Is Python a good programming language? Is Java? Is C? I think it&#x27;s a similar class of problem. You can benchmark rudimentary things like execution speed similar to how you can benchmark tokens&#x2F;second, but these metrics don&#x27;t tell the whole story.
    7. jdthedisciple · · focus · HN ↗
      Thank you. These seem to reasonably match my experience.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.