‹ BackHN Continuity

Thread

MiMo v2.6

1130 points · 483 comments · volf_

  1. rao-v · · focus · HN ↗
    I know we have strong views on what a truly open model is (open weights, open training data, open training code etc.) but I really like how transparent they’ve been about the training of this model.

    The realtime dashboard they shared during training (<a href="https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;rl&#x2F;" rel="nofollow">https:&#x2F;&#x2F;mimo.xiaomi.com&#x2F;rl&#x2F;) was an incredible learning and teaching tool for me, and they’ve been unusually comprehensive in sharing details about their methodology (check out that tech report - it&#x27;s got lots of clever behind the scene tricks like Google or Deepseek writeups) and benchmark scores (even the stuff they didn’t do well on).

    If you’re releasing an open model going forward, please consider offering the community more of this transparency!

    1. bicepjai · · focus · HN ↗
      I was absolutely mind blown when I saw how they were publishing that training dashboard while US models publish 100s of pages of reports (just provide a &quot;copy as MD&quot; button, folks, in the future). I was thinking about doing something similar but did not know how to show it, and this is a perfect example for someone who wants to show whatever they are training, for me it was local training on a consumer GPU.

      My dream is to see this like a dashboard for a model trained across distributed machines, like Bitcoin mining, where minted coins are given to people whose machines were used for training. I don&#x27;t know if they are worth it, but bragging rights alone, like a tag they can put on a website or social media, will be good enough for me.

      1. BodyCulture · · focus · HN ↗
        The distributed training collective is a good idea, let’s discuss details:

        A) How to prevent malicious injection of bad training data?

        B) How to handle copyright violations, will participants be responsible and will they have to pay the creators?

        C) Can this create an income stream for content creators and how to avoid abuse, eg feeding with AI content?

        So many more, but let’s focus on these before we break things fast because we didn’t think about them.

        1. PaulRobinson · · focus · HN ↗
          I think all of these can be addressed by being able to understand which machines contributed what to the model. I know that&#x27;s messy, but if I can state &quot;this information has come from Alice and the consensus is that it&#x27;s good and right as it aligns with other information from Bob and others, meanwhile this information came from Mallory and stands out as being incongruent with the rest of the information I have in the model&quot;, we can identify malicious violations. If we&#x27;re then able to state either Alice or Bob is one of the copyright holders, and one of them gets paid a little more, the other a little less as the confirmation agent, well, the economics of all of this changes a little.

          At the moment we have Annas Archive being paid by frontier labs and rare&#x2F;second hand books being destroyed in order to support the training regime. If instead we could just pay the publishers and they could distribute royalties to authors...

        2. dhx · · focus · HN ↗
          For (C) -- I think it could additional create jobs funded by government, philanthropic and other private institutions. For example, a government funded museum may already be participating in Wikimedia GLAM projects (e.g. uploading historical images to Wikimedia Commons with complete metadata). Perhaps this type of open source contribution may increase if organisations realise their mission can be better accomplished by contributing this same open data into LLMs, in addition to Wikimedia Commons. If the museum&#x27;s mission is to educate the public on the history of life in ACMEville, having LLMs be able to provide historical information and images to a prompt of &quot;What is the history of ACMEville?&quot; may be a good pursuit.

          I&#x27;m sceptical though whether use of LLMs would encourage creation of data that doesn&#x27;t already exist. For example, if you ask an LLM &quot;What are the top 100 most prevalent flora endemic to ACME National Park&quot;, this data may not currently exist _at all_, and to collect, would require paying botanists to do an extensive field survey. If no one has done this work yet--why? Is it relevant to the scientific community, to making government decisions, etc, or just an obscure academic curiosity. There are certainly some journal articles on _other_ national parks describing some of their common endemic flora, but perhaps there was a reason for this. Such as a scientist funded by a one-off government program trying to determine how to preserve or even create habitat for a specific endangered species.

          Consider for the prompt of: &quot;What are the top 100 most prevalent flora endemic to ACME National Park&quot;

          An LLM may reply: &quot;I couldn&#x27;t find any journal article or other prior work that may answer this question. Typically such survey field work may cost $X to complete, require expert botanists, and take 6-12 months to complete. Let me know if you want further information on how to find and select a company to conduct such a botanical field survey.&quot;

          Would this type of LLM response grow the industry of botanical field surveys, or do nothing, perhaps because anyone likely to fund botanical field surveys is already doing so regardless of whatever is happening with AI.

          1. bicepjai · · focus · HN ↗
            Sorry. This seems like llm generated nonsense :)
        3. bicepjai · · focus · HN ↗
          Thank for starting a thinking thread. The way I thought about this was everything is preset, model, data (prepared, cleaned) and we cannot prevent duplicate training but find a way to get the most recent weights is of now and share it when possible to synchronize; most recent weights always available on torrent. Now addressing your questions A) how do I trust the weights passes along is when we can use the crypto proof of stake for which we need some kind of coin I suppose. B) dataset is prepared so, it’s open. I don’t know if that helps or not C) What has this to do with content creators ? You mean when content creators use the model, they get coins ? I have not thought about that angle.

          This can be a fun conversation and discussion.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.