‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. rao-v · · focus · HN ↗
    I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.

    The realtime dashboard they shared during training (<a href="https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;rl&#x2F;" rel="nofollow">https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;rl&#x2F;) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it&#x27;s got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).

    If you’re releasing an open model going forward, please consider offering the community more of this transparency!

    1. MangoCoffee · · focus · HN ↗
      maybe this is why Dario want to slow down AI development and all the big AI labs in the USA is singing the same song.

      whey they all singing the same tune. it make me question what is their real motives.

      they are afraid of Chinese good enough LLM model killing their margin. we already have story about US companies switch some task to use cheaper Chinese model hosted on Neoclouds.

      1. aeyes · · focus · HN ↗
        The reason is money. They want regulation to make it harder for new competitors and competitors from other countries.

        They invested billions into training the models but there is no competitive advantage, we see that within a couple of months everyone catches up. There is no way to profitability unless they get some policies to shields them against competitors that can&#x27;t comply with the regulatory requirements.

        That is also why there are things like Claude, Codex and Cursor. They are trying hard to build a customer relationship with a higher switching cost that hopefully sticks.

        But the problem is that the AI buildout has become a large percentage of GDP. So obviously the government wants to keep it going because these companies are pumping enormous amounts of money into the economy.

        1. vlovich123 · · focus · HN ↗
          Unless something has shifted, “everyone catches up” is because these bleeding edge models are distilled. You don’t see this happening with other European and US labs and the problem isn’t something being ignored. I’m not convinced this pattern will continue indefinitely.
          1. awad · · focus · HN ↗
            Why is it OK to train on the collective IP of humanity and call it fair use but then call the next batch distilled with negative connotations?
            1. userbinator · · focus · HN ↗
              This is why Imaginary Property is an illusion, as everything is a derivative work, and AI is going to make that fact even clearer.
              1. eru · · focus · HN ↗
                That&#x27;s not true for literally everything.

                When eg I snap a picture of my dog, that&#x27;s not derived from anything. But I still get intellectual property rights for the photograph.

            2. floam · · focus · HN ↗
              I don’t follow. Fair use is a copyright defense, and nobody is suggesting distillation attacks are just a copyright violation are they?

              Aren’t they alleging these other companies directly entered into a contract and violated the terms, and in cases where question, answer pairs were obtained without such agreement, it was accomplished by outright wire fraud or theft?

              1. kkotak · · focus · HN ↗
                Are you suggesting that worldwide copyright violation is more acceptable than contract breach between companies?
                1. floam · · focus · HN ↗

                  [dead]

                2. [deleted] · · focus · HN ↗

                  [deleted]

            3. vlovich123 · · focus · HN ↗
              I did no such moral claim. I just noted that the foundation labs are working on technical hurdles to thwart distillation efforts and the cost and quality of Chinese models isn’t likely to keep up with the 6 month lag time everyone has assumed.
              1. awad · · focus · HN ↗
                Fair enough, apologies for reading in to it that which you did not mean.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.