Been playing around with both of these since last night. So far, when enabling reasoning (which is binary on/off), it seems to me like Flash either is less "token efficient" or just likes to think more, or I'm doing something else wrong, because most prompts I send to both, Flash reasons more and for longer than Pro, which is the opposite of my expectations.
Any lower (or broken) quantization could do that, not what I'm talking about though. They work fine for their size, as far as I can tell. Just surprised the Flash would reason for longer than the Pro.
stymaar · · focus · HN ↗
Pro [2]:, 1.02T total / 42B activated parameters
[1]: <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL" rel="nofollow">https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
[2]: <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL" rel="nofollow">https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
embedding-shape · · focus · HN ↗
Are others seeing the same thing?
the_duke · · focus · HN ↗
This is also true for Deepseek 4(.1) .
embedding-shape · · focus · HN ↗