‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. user43928 · · focus · HN ↗
    I don't trust any of the benchmarks where Opus 5 surpasses Astra or Fable 5.1.

    Maybe Terminal Bench 4.0 and ExploitGym are reasonable.

    Terminal Bench 4.0

      GPT 6 Astra             59.6
      Claude Fable 5.1        55.1
      Claude Opus 5           49.0
      MiMo-V2.6-Pro           34.9
      MiMo-V2.6-Flash         28.8
      DeepSeek V4.1 Flash     26.8
      MiMo-V2.5-Pro            1.5
    
    ExploitGym

      GPT 6 Astra             42.4
      Claude Fable 5.1        30.4
      Claude Opus 5           22.1
      MiMo-V2.6-Pro           17.8
      MiMo-V2.6-Flash          6.0
      MiMo-V2.5-Pro            0.1
    
    DeepSWE v1.1

      DeepSeek V4.1 Flash     74.2
      Claude Opus 5           74.0
      GPT 6 Astra             74.0
      MiMo-V2.6-Pro           71.9
      Claude Fable 5          70.0
      MiMo-V2.6-Flash         67.9
      MiMo-V2.5-Pro           19.0
    1. dom96 · · focus · HN ↗
      Why not? In my own benchmark Opus 5 does in fact come out on top[1]

      1 - <a href="https:&#x2F;&#x2F;bench.killswitch-lang.org&#x2F;" rel="nofollow">https:&#x2F;&#x2F;bench.killswitch-lang.org&#x2F;

      1. user43928 · · focus · HN ↗
        Good question, maybe I am underestimating it based on its absolutely horrible writing style.
        1. cosmojg · · focus · HN ↗
          Yeah, for better or worse, writing style is practically uncorrelated with agentic performance, which is all the rage right now and the thing that most popular benchmarks currently prioritize.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.