Been playing around with both of these since last night. So far, when enabling reasoning (which is binary on/off), it seems to me like Flash either is less "token efficient" or just likes to think more, or I'm doing something else wrong, because most prompts I send to both, Flash reasons more and for longer than Pro, which is the opposite of my expectations.
Tried both <a href="https://token-plan-ams.xiaomimimo.com/v1" rel="nofollow">https://token-plan-ams.xiaomimimo.com/v1 and <a href="https://api.xiaomimimo.com/v1" rel="nofollow">https://api.xiaomimimo.com/v1, hoping they don't run any quantized variants there, but who knows.
stymaar · · focus · HN ↗
Pro [2]:, 1.02T total / 42B activated parameters
[1]: <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL" rel="nofollow">https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
[2]: <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL" rel="nofollow">https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
embedding-shape · · focus · HN ↗
Are others seeing the same thing?
alfiedotwtf · · focus · HN ↗
embedding-shape · · focus · HN ↗