‹ BackHN Continuity

Thread

MiMo-v2.6-Pro: Intelligence, Performance and Price Analysis

166 points · 68 comments · theanonymousone

  1. egeres · · focus · HN ↗
    It feels suspicious that MiMo-V2.6 Pro gets 46 in de index while DeepSeek-V4.1 (<a href="https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;deepseek-v4-1-flash" rel="nofollow">https:&#x2F;&#x2F;artificialanalysis.ai&#x2F;models&#x2F;deepseek-v4-1-flash) gets 39. According to the appendix at the bottom of <a href="https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;mimo-v2-6" rel="nofollow">https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;mimo-v2-6 the deepseek model sometimes surpasses mimo and it&#x27;s not so far behind in capabilities. A week ago opus 5 appeared 1 points ahead of fable 5 despite fable being a much smarter model (this has been corrected already)
    1. SyneRyder · · focus · HN ↗
      The main AA benchmark keeps changing, and had to be radically changed when Astra came out and showed zero improvement over GPT 5.6 Sol in their benchmark. Opus 5 is still 1 point ahead of Fable 5.0 on the index, if you manually add Fable 5.0 back into the list, so it hasn&#x27;t actually been &quot;corrected&quot;. It&#x27;s only Fable 5.1 that is shown as ahead of Opus 5.

      The AA benchmark is a weighted average of other benchmarks and some internal ones. I think the difficult part is finding benchmarks that reflect your own use of the models.

      1. seahorseemoji · · focus · HN ↗
        The way Artificial Analysis keeps changing their weights feels kind of like deciding who the winner should be and making the weights reflect that. They’ve been changing their weights to add more weight to improved long-running agentic capabilities, but doing so means they’re reducing the relative importance of world knowledge and of writing ability.

        I’ll grant that maybe world knowledge isn’t that important for these models. But writing ability is important for human understanding, and I think the weird turns of phrase and word choices reflect the labs’ underweighting of the importance of human understanding.

        1. sipjca · · focus · HN ↗
          I mean artificial is in their name....
        2. polishmijajca · · focus · HN ↗
          How does someone objectively quantify writing ability?
      2. yt1998 · · focus · HN ↗

        [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.