I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.
The realtime dashboard they shared during training (<a href="https://mimo.xiaomi.com/rl/" rel="nofollow">https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).
If you’re releasing an open model going forward, please consider offering the community more of this transparency!
maybe this is why Dario want to slow down AI development and all the big AI labs in the USA is singing the same song.
whey they all singing the same tune. it make me question what is their real motives.
they are afraid of Chinese good enough LLM model killing their margin. we already have story about US companies switch some task to use cheaper Chinese model hosted on Neoclouds.
Please explain how putting an upper bound on how good the strongest models can be prevents cheaper less strong models from catching up, rather than enabling it. I do not understand this argument at all.
The general idea is that Anthropic/OpenAI is pushing this narrative as an attempt at "Regulatory Capture"[1] which would allow them to make it prohibitively expensive for anyone but them to enter the market thus stifling competition.
Local LLM's are hit worse. Its about 6k for 5090 or 15k for an RTX 6000 and the Mac Ultra 256 is upwards of 12k.
Sure if you already have hardware, you can frankenbuild a system - but even the "affordable" dev stations of the DGX Sparks went from 3500 to 5k and upwards of 8k depending on vendor.
All the meanwhile, OpenAI pushed Luna 6 which is crazy cheap suggesting they have flash models to compete with Chinese models.
I just hope we see more open weights. Nvidia has NEMO but their license doesn't allow NEMO to be re-used on say, Apple or ROCm - Nvidia hardware only. Big ol MEH
rao-v · · focus · HN ↗
The realtime dashboard they shared during training (<a href="https://mimo.xiaomi.com/rl/" rel="nofollow">https://mimo.xiaomi.com/rl/) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it's got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).
If you’re releasing an open model going forward, please consider offering the community more of this transparency!
MangoCoffee · · focus · HN ↗
whey they all singing the same tune. it make me question what is their real motives.
they are afraid of Chinese good enough LLM model killing their margin. we already have story about US companies switch some task to use cheaper Chinese model hosted on Neoclouds.
jwolfe · · focus · HN ↗
rbjorklin · · focus · HN ↗
* 1: <a href="https://en.wikipedia.org/wiki/Regulatory_capture" rel="nofollow">https://en.wikipedia.org/wiki/Regulatory_capture
supernovae · · focus · HN ↗
Local LLM's are hit worse. Its about 6k for 5090 or 15k for an RTX 6000 and the Mac Ultra 256 is upwards of 12k.
Sure if you already have hardware, you can frankenbuild a system - but even the "affordable" dev stations of the DGX Sparks went from 3500 to 5k and upwards of 8k depending on vendor.
All the meanwhile, OpenAI pushed Luna 6 which is crazy cheap suggesting they have flash models to compete with Chinese models.
I just hope we see more open weights. Nvidia has NEMO but their license doesn't allow NEMO to be re-used on say, Apple or ROCm - Nvidia hardware only. Big ol MEH