‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. user43928 · · focus · HN ↗
    I don't trust any of the benchmarks where Opus 5 surpasses Astra or Fable 5.1.

    Maybe Terminal Bench 4.0 and ExploitGym are reasonable.

    Terminal Bench 4.0

      GPT 6 Astra             59.6
      Claude Fable 5.1        55.1
      Claude Opus 5           49.0
      MiMo-V2.6-Pro           34.9
      MiMo-V2.6-Flash         28.8
      DeepSeek V4.1 Flash     26.8
      MiMo-V2.5-Pro            1.5
    
    ExploitGym

      GPT 6 Astra             42.4
      Claude Fable 5.1        30.4
      Claude Opus 5           22.1
      MiMo-V2.6-Pro           17.8
      MiMo-V2.6-Flash          6.0
      MiMo-V2.5-Pro            0.1
    
    DeepSWE v1.1

      DeepSeek V4.1 Flash     74.2
      Claude Opus 5           74.0
      GPT 6 Astra             74.0
      MiMo-V2.6-Pro           71.9
      Claude Fable 5          70.0
      MiMo-V2.6-Flash         67.9
      MiMo-V2.5-Pro           19.0
    1. novaleaf · · focus · HN ↗
      can you recommend any benchmark websites that show up-to-date details like this?

      TerminaBench, DeepSwe sites are out of date.

      1. UnfitFootprint · · focus · HN ↗
        Yeah it’s a shame a lot of these benchmarks are behind. My favourite was ‘SlopCodeBench’ [1] as I’m most interested in ai reinforcing its own bad decisions, but it’s not even up to current gen oai

        1: <a href="https:&#x2F;&#x2F;www.scbench.ai&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.scbench.ai&#x2F;

      2. 1899-12-30 · · focus · HN ↗
        <a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;evaluations&#x2F;terminalbench-4-0" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;evaluations&#x2F;terminalbench-4-0
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.