‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. rao-v · · focus · HN ↗
    I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.

    The realtime dashboard they shared during training (<a href="https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;rl&#x2F;" rel="nofollow">https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;rl&#x2F;) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it&#x27;s got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).

    If you’re releasing an open model going forward, please consider offering the community more of this transparency!

    1. earthnail · · focus · HN ↗
      Thanks so much for sharing this. As someone who mostly watches from the sideline, can you share what you can see in this dashboard that someone like me can&#x27;t see? Is it the metrics themselves that they measure (the metrics tab is absurdly detailed), something in the notices, or something else I missed?
      1. verdverm · · focus · HN ↗
        the existence, who else has a live dashboard for the RL late-training?
      2. tancop · · focus · HN ↗
        The best thing they did is being open about all the setbacks they had to deal with. They logged every restart with a reason, talked about dropping a cyber dataset after it degraded coding benchmarks. Also published real time training loss, benchmark scores after every checkpoint and running cost estimates.

        Really the only thing missing was dataset descriptions, the dashboard only had random IDs like &quot;dataset-zrso&quot;. I guess it&#x27;s their lawyers fault.

        1. dhx · · focus · HN ↗
          100% agreed

          It&#x27;d be great to see a description of even just a subset of training datasets. It feels very much under-reported how much expense is worth investing in preparing and selecting training datasets versus just using masses of random quality unprepared training data. This dashboard appears to be good though in showing the limits quickly reached when throwing parameters and compute at the problem.

          For example, if they were to train on Wikipedia dumps, do they consider every article to be the same quality across each language, or have they done more work beyond Wikipedia&#x27;s own article quality ratings to make training decisions such as &quot;Ignore cebwiki it&#x27;s machine-generated spam&quot; and &quot;Treat dewiki articles with coordinates within Germany as being higher quality (weight it higher) than their equivalent enwiki articles&quot;.

          And let&#x27;s say one of the datasets is all the source code of packages in the Gentoo package repository. Not every software package is a good example of how to write code. You perhaps wouldn&#x27;t want to train your LLM on 1990s era PHP web application source code as an example of how to write code in 2026. Instead, you&#x27;d possibly want to use such PHP web application source code as a negative training example of what _not_ to write. But when training an LLM to detect software bugs, maybe outdated PHP source code is good for training.

          Similarly for translation, perhaps UN treaty documents translated into 4+ languages are good translation examples because of high accuracy needed, professional translators being used, and bigger budgets. However this training data would perhaps be a negative training example towards translating chat messages, movie subtitles, etc because it doesn&#x27;t use everyday slang and could result in output of nonsense such as &quot;Pending Your Excellency&#x27;s response, please accept, Your Excellency, my sincere greetings.&quot; for a prompt asking to write a birthday card for a child.

          Preparing training data and deciding how to best use it for training I assume would be the largest expense (cost of labour -- mostly expert labour too) and also the greatest opportunity in the future for LLMs to improve. It seems to me somewhat irrelevant if the dashboard indicates a compute expense of $1m or $5m if good training datasets (prepared by experts in their fields) cost $10m&#x2F;y to maintain. For example, hiring expert software developers to tag 1000&#x27;s of open source software packages according to their quality, on different metrics, such as human readability, performance optimisation with choice of algorithms, reasonable trade-off between coherence and coupling in the software architecture, currency with state of the art programming trends&#x2F;preferred dependencies&#x2F;operating system APIs, etc. And keeping that metadata continually updated rather than a rapidly obsolete once off tagging project completed in 2005.

      3. rao-v · · focus · HN ↗
        I might turn this into a blogpost if folks are interested, but my god there is so much clever info in that dashboard.

        Here is one really neat bit:

        A cutting edge training idea (for agents, it&#x27;s been used elsewhere for ages) is on-policy RL, basically, it&#x27;s not enough to say &quot;here is an end to end agentic sequence (including tool calls etc.) that is perfect&quot; you want to say &quot;here is a sequence you might actually have generated that turns out to be correct&quot;.

        Basically, it&#x27;s more training efficient to improve models with small tweaks to do more of the right thing they are already doing sometimes than from some perfect oracular &quot;this is the way&quot; answer.

        (if you&#x27;ve ever tried to teach humans new skills, you’ve probably noticed this too!)

        When you do that, you care about how far the model you are updating (improving) has deviated from the one being used to generate rollouts (agentic rollouts for hard problems can take hours with lots of tool calls, so you can&#x27;t keep redeploying every slight improvement).

        Lo and behold, the dashboard literally has:

        partial&#x2F;avg_staleness (likely the measure of how many micro iterations the &quot;generate answers&quot; model is behind the &quot;improving based on the occasional right answer&quot; model)

        train_infer_diff&#x2F;new_infer&#x2F;kl (a more direct KL divergence based way of measuring how differently the two models generate tokens)

        How cool is that?!

        And don&#x27;t get me started on the clever ideas hiding behind dynsam&#x2F;avg@n ...

        1. dgellow · · focus · HN ↗
          Please do
        2. jeffmcjunkin · · focus · HN ↗
          I&#x27;d read the heck out of that blogpost. You have my interest.
        3. pimeys · · focus · HN ↗
          I would really enjoy that blog post.
        4. oceansweep · · focus · HN ↗
          Please do!
        5. armas · · focus · HN ↗
          yes I&#x27;m interested. please consider writing this
        6. k9294 · · focus · HN ↗
          +1 waiting for the blog post!
        7. Bluestein · · focus · HN ↗
          Seconded.-
        8. derpyzza · · focus · HN ↗
          definitely do!
        9. handfuloflight · · focus · HN ↗
          Write and we shall read.
        10. lemontheme · · focus · HN ↗
          Hold on, isn&#x27;t that just standard practice for post-training LLMs for agentic use? Give task, generate n rollouts, grade rollouts (either at termination or after each tool call)? Or is the difference that the rollouts are generated ahead of time and then graded? (Of course, then it&#x27;s not really on-policy.)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.