‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. stymaar · · focus · HN ↗
    Flash[1]: 309B total / 15B activated parameters

    Pro [2]:, 1.02T total / 42B activated parameters

    [1]: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Flash-RL" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Flash-RL

    [2]: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Pro-RL" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Pro-RL

    1. verdverm · · focus · HN ↗
      curious why the HF pill (on the right) always has inaccurate values
      1. stymaar · · focus · HN ↗
        I noticed the same, and I wonder as well.
        1. verdverm · · focus · HN ↗
          I suspect they are calculating something in the weights or config, I see it pretty consistently with quants
      2. bopbop9876 · · focus · HN ↗
        I believe it&#x27;s because this model is natively fp8 (for the most part), and that display struggles native quants.
      3. wren6991 · · focus · HN ↗
        Packed 4-bit weights are often identified as u8 byte arrays in the safetensors metadata, so the HF UI counts them as half a weight each.
        1. verdverm · · focus · HN ↗
          I would think there is sufficient information in the various config files to work this out, based what I&#x27;ve seen in my own quant artifacts.
          1. wren6991 · · focus · HN ↗
            Yeah, I agree it&#x27;s probably fixable, but I think a naive interpretation of the metadata is probably the source of this bug (which HF has had for as long as I can remember)
    2. verdverm · · focus · HN ↗
      There&#x27;s also a Qwen 3.5 9B distill

      <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Distill-Qwen-9B" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;XiaomiMiMo&#x2F;MiMo-V2.6-Distill-Qwen-9B

      1. gandreani · · focus · HN ↗
        Those this mean they&#x27;ve fine-tuned this Qwen 3.5 9B on output from the V2.6 model?
        1. mydreamof · · focus · HN ↗
          It is a 9B agentic model developed by Xiaomi MiMo through supervised fine-tuning of Qwen3.5-9B on MiMo-generated data
        2. simonedepertis · · focus · HN ↗

          [dead]

    3. segmondy · · focus · HN ↗
      more like 500B in FP8
    4. embedding-shape · · focus · HN ↗
      Been playing around with both of these since last night. So far, when enabling reasoning (which is binary on&#x2F;off), it seems to me like Flash either is less &quot;token efficient&quot; or just likes to think more, or I&#x27;m doing something else wrong, because most prompts I send to both, Flash reasons more and for longer than Pro, which is the opposite of my expectations.

      Are others seeing the same thing?

      1. alfiedotwtf · · focus · HN ↗
        Maybe it’s the quant you’re using?
        1. embedding-shape · · focus · HN ↗
          Tried both <a href="https:&#x2F;&#x2F;token-plan-ams.xiaomimimo.com&#x2F;v1" rel="nofollow">https:&#x2F;&#x2F;token-plan-ams.xiaomimimo.com&#x2F;v1 and <a href="https:&#x2F;&#x2F;api.xiaomimimo.com&#x2F;v1" rel="nofollow">https:&#x2F;&#x2F;api.xiaomimimo.com&#x2F;v1, hoping they don&#x27;t run any quantized variants there, but who knows.
      2. the_duke · · focus · HN ↗
        Most of the Chinese models get into long complicated thinking loops before they accomplish something more complicated.

        This is also true for Deepseek 4(.1) .

        1. embedding-shape · · focus · HN ↗
          Any lower (or broken) quantization could do that, not what I&#x27;m talking about though. They work fine for their size, as far as I can tell. Just surprised the Flash would reason for longer than the Pro.
        2. FredFS456 · · focus · HN ↗
          I prefer GLM 5.3 (Flash or not) over Deepseek 4.1 Flash because the GLM models are significantly less verbose.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.