‹ BackHN Continuity

Thread

MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis

166 points · 68 comments · theanonymousone

  1. egeres · · focus · HN ↗
    It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 (<a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;deepseek-v4-1-flash" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;deepseek-v4-1-flash) gets 39. According to the appendix at the bottom of <a href="https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;mimo-v2-6" rel="nofollow">https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;mimo-v2-6 the deepseek model sometimes surpasses mimo and it&#x27;s not so far behind in capabilities. A week ago opus 5 appeared 1 points ahead of fable 5 despite fable being a much smarter model (this has been corrected already)
    1. SyneRyder · · focus · HN ↗
      The main AA benchmark keeps changing, and had to be radically changed when Astra came out and showed zero improvement over GPT 5.6 Sol in their benchmark. Opus 5 is still 1 point ahead of Fable 5.0 on the index, if you manually add Fable 5.0 back into the list, so it hasn&#x27;t actually been &quot;corrected&quot;. It&#x27;s only Fable 5.1 that is shown as ahead of Opus 5.

      The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.

      1. yt1998 · · focus · HN ↗

        [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.