Flash[1]: 309B total / 15B activated parametersPro [2]:, 1.02T total / 42B activated parameters[1]: <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL" rel="nofollow">https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL[2]: <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL" rel="nofollow">https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
I suspect they are calculating something in the weights or config, I see it pretty consistently with quants
I believe it's because this model is natively fp8 (for the most part), and that display struggles native quants.
Packed 4-bit weights are often identified as u8 byte arrays in the safetensors metadata, so the HF UI counts them as half a weight each.
I would think there is sufficient information in the various config files to work this out, based what I've seen in my own quant artifacts.
Yeah, I agree it's probably fixable, but I think a naive interpretation of the metadata is probably the source of this bug (which HF has had for as long as I can remember)
stymaar · · focus · HN ↗
Pro [2]:, 1.02T total / 42B activated parameters
[1]: <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL" rel="nofollow">https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Flash-RL
[2]: <a href="https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL" rel="nofollow">https://huggingface.co/XiaomiMiMo/MiMo-V2.6-Pro-RL
verdverm · · focus · HN ↗
stymaar · · focus · HN ↗
verdverm · · focus · HN ↗
bopbop9876 · · focus · HN ↗
wren6991 · · focus · HN ↗
verdverm · · focus · HN ↗
wren6991 · · focus · HN ↗