‹ BackHN Continuity

Thread

Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

310 points · 140 comments · leonickson

  1. dghlsakjg · · focus · HN ↗
    I know everyone wants to crap all over these setups that are impractical, but this is how progress happens.

    People will keep plugging away at this and figure out how to avoid wearing the hard drive, how to make it run faster, custom hardware buses etc.

    Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.

    1. pizza234 · · focus · HN ↗
      > but this is how progress happens.

      This is progress in the same way that a man climbing a tree is making progress toward reaching the moon.

      This project is essentially the MoE-of-the day, with some platform-related optimizations.

      > Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.

      That won't happen. Projects like this just give the illusion that that will be possible.

      1. trymas · · focus · HN ↗
        it’s equivalent take to laugh at first transformers 9 years ago, because they were shit and hardware requirements were immense.

        IMHO it’s a matter of time until we (consumers) will get the hardware (maybe coupled maybe even more novel techniques). Though I expect it will take another 10 years or more.

        1. mv4 · · focus · HN ↗
          The frontier labs will do their best to prevent this from happening. Their financial model won't work if people start running open-weight models on their local hardware.

          This is why banning Chinese open-weight AI models is a major policy debate in Washington. The labs can't survive log-term without subsidies, and a ban can act as a subsidy.

          1. trymas · · focus · HN ↗
            I concur, that this could be a possiblity and it definitely seems now.

            Though my bet would be, if USA will go ultra protectionist in this regard - in 10-20 years most world will run Chinese LLMs and hardware for this purpose.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.