‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. stymaar · · focus · HN ↗
    Flash[1]: 309B total / 15B activated parameters

    Pro [2]:, 1.02T total / 42B activated parameters

    [1]: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Flash-RL" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Flash-RL

    [2]: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Pro-RL" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Pro-RL

    1. embedding-shape · · focus · HN ↗
      Been playing around with both of these since last night. So far, when enabling reasoning (which is binary on&#x2F;off), it seems to me like Flash either is less &quot;token efficient&quot; or just likes to think more, or I&#x27;m doing something else wrong, because most prompts I send to both, Flash reasons more and for longer than Pro, which is the opposite of my expectations.

      Are others seeing the same thing?

      1. alfiedotwtf · · focus · HN ↗
        Maybe it’s the quant you’re using?
        1. embedding-shape · · focus · HN ↗
          Tried both <a href="https:&#x2F;&#x2F;token-plan-ams.xiaomimimo.com&#x2F;v1" rel="nofollow">https:&#x2F;&#x2F;token-plan-ams.xiaomimimo.com&#x2F;v1 and <a href="https:&#x2F;&#x2F;api.xiaomimimo.com&#x2F;v1" rel="nofollow">https:&#x2F;&#x2F;api.xiaomimimo.com&#x2F;v1, hoping they don&#x27;t run any quantized variants there, but who knows.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.