‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. rao-v · · focus · HN ↗
    I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.

    The realtime dashboard they shared during training (<a href="https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;rl&#x2F;" rel="nofollow">https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;rl&#x2F;) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it&#x27;s got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).

    If you’re releasing an open model going forward, please consider offering the community more of this transparency!

    1. MangoCoffee · · focus · HN ↗
      maybe this is why Dario want to slow down AI development and all the big AI labs in the USA is singing the same song.

      whey they all singing the same tune. it make me question what is their real motives.

      they are afraid of Chinese good enough LLM model killing their margin. we already have story about US companies switch some task to use cheaper Chinese model hosted on Neoclouds.

      1. aeyes · · focus · HN ↗
        The reason is money. They want regulation to make it harder for new competitors and competitors from other countries.

        They invested billions into training the models but there is no competitive advantage, we see that within a couple of months everyone catches up. There is no way to profitability unless they get some policies to shields them against competitors that can&#x27;t comply with the regulatory requirements.

        That is also why there are things like Claude, Codex and Cursor. They are trying hard to build a customer relationship with a higher switching cost that hopefully sticks.

        But the problem is that the AI buildout has become a large percentage of GDP. So obviously the government wants to keep it going because these companies are pumping enormous amounts of money into the economy.

        1. lelanthran · · focus · HN ↗
          &gt; But the problem is that the AI buildout has become a large percentage of GDP. So obviously the government wants to keep it going because these companies are pumping enormous amounts of money into the economy.

          They are pumping enormous amounts of money into each other. Hardly any of that is making its way to people, it&#x27;s all going to highly automated construction and to energy use.

          Seriously, how many jobs did the $1t in venture capital fund?

          1. gunalx · · focus · HN ↗
            If I pay you 100$ for mowing my lawn, and you me for yours. Technically the GDP increased with 200$.
            1. genxy · · focus · HN ↗
              How about I draw you a picture instead. Mowing a lawn is a priceable service.
            2. cosmojg · · focus · HN ↗
              And, in this case, the dollar-amount increase in GDP serves as a virtual quantitative proxy for the increase in mowed lawns (and the value thereof). In other words, the participants in this economy are collectively ~$200 richer with their mowed lawns than they were without them.
              1. AlotOfReading · · focus · HN ↗
                This is a thinly disguised broken window parable.

                If everyone goes around mowing lawns for each other, the economy is richer in lawn mowing at the expense of all the other things that would have been funded had everyone mowed their own lawns and purchased different services instead.

                1. aytigra · · focus · HN ↗
                  I am confused with this, if &quot;everyone mowed their own lawns&quot; then the net result will be exactly the same, everyone will be busy the same and not poorer, just without money movement.
                  1. saidnooneever · · focus · HN ↗
                    look at the broken window parable as he mentioned it might help understand the rest of his comment
                    1. tacitusarc · · focus · HN ↗
                      Broken window is different from the mowing lawns hypothetical
                    2. adammarples · · focus · HN ↗
                      This is not the same. If everyone wants mowed lawns, and everyone is busy working on that, there is no opportunity cost, everyone is working on their top priorities. The broken window fallacy is a fallacy because the headline gdp figure doesn&#x27;t account for the destruction of the window which cancels out the benefit. In the grass mowing analogy nothing has been destroyed, useful and priority work has been done all around.
              2. socialcommenter · · focus · HN ↗
                If the pricing is fair and at arms&#x27; length. What&#x27;s happening in reality is as if they are mowing each others&#x27; lawns at wink wink nudge nudge $1000. Not a good proxy for actual value created.
                1. eru · · focus · HN ↗
                  In the real world, you have to pay taxes. So people are incentivized to claim less value for the lawns mowed, or even just do it themselves, instead of benefiting from the division of labour.
                  1. j1mmyrustl3r · · focus · HN ↗

                    [dead]

                  2. socialcommenter · · focus · HN ↗
                    Person A has leverage, and every $1000 sale makes his share price $10000 higher, more than compensating for the $100 in taxes.

                    Person B owns shares in Person A.

                    &gt; eru

                    Tolkien fan?

                    1. eru · · focus · HN ↗
                      Leverage doesn&#x27;t work that way. (If it were so easy, it would load up my investment portfolio with a lot more leverage than I currently do.)

                      Tolkien is great, yes.

              3. Naracion · · focus · HN ↗
                But also importantly the government of the residents&#x27; country is about 39% ($78) richer, if say the participants are honest in reporting this and the country is the UK and the participants are people like you and me in the tech industry who frequent HN and would think to do something like this.
            3. Wowfunhappy · · focus · HN ↗
              Well, yes, because both of your lawns got mowed!

              Value was created!

              1. [deleted] · · focus · HN ↗

                [deleted]

            4. [deleted] · · focus · HN ↗

              [deleted]

          2. eru · · focus · HN ↗
            &gt; Hardly any of that is making its way to people, it&#x27;s all going to highly automated construction and to energy use.

            How do we know that? How automated is the construction really?

            In any case, the Fed and other central banks can print as much money as they want in order to hit any aggregate spending or inflation target they have for the economy.

        2. jamienk · · focus · HN ↗
          I’m still not at all sure about the “billions” invested claim. How much of that is cloud running the models? How much is pre and post training (which may or may not be part of what we’d want to include in accounting). Etc. Does anyone have links to good reporting about this: not blind recitations of numbers, but analysis and thought mixes with investigation?
        3. vlovich123 · · focus · HN ↗
          Unless something has shifted, “everyone catches up” is because these bleeding edge models are distilled. You don’t see this happening with other European and US labs and the problem isn’t something being ignored. I’m not convinced this pattern will continue indefinitely.
          1. awad · · focus · HN ↗
            Why is it OK to train on the collective IP of humanity and call it fair use but then call the next batch distilled with negative connotations?
            1. userbinator · · focus · HN ↗
              This is why Imaginary Property is an illusion, as everything is a derivative work, and AI is going to make that fact even clearer.
              1. eru · · focus · HN ↗
                That&#x27;s not true for literally everything.

                When eg I snap a picture of my dog, that&#x27;s not derived from anything. But I still get intellectual property rights for the photograph.

            2. floam · · focus · HN ↗
              I don’t follow. Fair use is a copyright defense, and nobody is suggesting distillation attacks are just a copyright violation are they?

              Aren’t they alleging these other companies directly entered into a contract and violated the terms, and in cases where question, answer pairs were obtained without such agreement, it was accomplished by outright wire fraud or theft?

              1. kkotak · · focus · HN ↗
                Are you suggesting that worldwide copyright violation is more acceptable than contract breach between companies?
                1. floam · · focus · HN ↗

                  [dead]

                2. [deleted] · · focus · HN ↗

                  [deleted]

            3. vlovich123 · · focus · HN ↗
              I did no such moral claim. I just noted that the foundation labs are working on technical hurdles to thwart distillation efforts and the cost and quality of Chinese models isn’t likely to keep up with the 6 month lag time everyone has assumed.
              1. awad · · focus · HN ↗
                Fair enough, apologies for reading in to it that which you did not mean.
          2. Balinares · · focus · HN ↗
            I don&#x27;t understand on what you are basing this reasoning? If a well educated workforce can produce Fable then why couldn&#x27;t a well educated workforce produce MiMo?

            Besides, the latter actually published and open-sourced its RL stack to make it reproducible, which would in fact make it more trustworthy than the models you are speculating were distilled.

        4. Shekelphile · · focus · HN ↗
          No chinese lab has caught up yet. They&#x27;ve tried to fake it by distilling and overfitting on benchmarks to make their models look better than they are, the &#x27;best&#x27; models available from chinese labs right now (GLM 5.3 and Kimi K3) fall apart completely when you try to do real work with them. K3 is especially embarrassing because it is larger than Mythos yet performs worse than opus 5 and 5.6 sol in benchmarks they haven&#x27;t been able to fake yet.
          1. kkotak · · focus · HN ↗
            In that case, Open AI and Anthropic have nothing to worry about.
        5. eru · · focus · HN ↗
          Well, if you are right, I just hope their protectionism will only affect the American market, and they leave us unAmericans free to get our models from wherever.
        6. barrenko · · focus · HN ↗
          On the other hand (OTOH), China is desperate to keep up and keeps pushing open models (rightfully so), as they understand how far ahead from everyone the US is, and that whoever gets this right first basically is going to become an alien compared to others.

          But even with all the open models the US is just insanely ahead in AI buildout and capital allocation (as usual).

          Is it a bubble? Is it like the race for the-first-to-the-nuclear bomb? Both?

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.