‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. stymaar · · focus · HN ↗
    Flash[1]: 309B total / 15B activated parameters

    Pro [2]:, 1.02T total / 42B activated parameters

    [1]: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Flash-RL" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Flash-RL

    [2]: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Pro-RL" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Pro-RL

    1. embedding-shape · · focus · HN ↗
      Been playing around with both of these since last night. So far, when enabling reasoning (which is binary on&#x2F;off), it seems to me like Flash either is less &quot;token efficient&quot; or just likes to think more, or I&#x27;m doing something else wrong, because most prompts I send to both, Flash reasons more and for longer than Pro, which is the opposite of my expectations.

      Are others seeing the same thing?

      1. the_duke · · focus · HN ↗
        Most of the Chinese models get into long complicated thinking loops before they accomplish something more complicated.

        This is also true for Deepseek 4(.1) .

        1. embedding-shape · · focus · HN ↗
          Any lower (or broken) quantization could do that, not what I&#x27;m talking about though. They work fine for their size, as far as I can tell. Just surprised the Flash would reason for longer than the Pro.
        2. FredFS456 · · focus · HN ↗
          I prefer GLM 5.3 (Flash or not) over Deepseek 4.1 Flash because the GLM models are significantly less verbose.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.