‹ BackHN Continuity

Thread

Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

310 points · 140 comments · leonickson

  1. dghlsakjg · · focus · HN ↗
    I know everyone wants to crap all over these setups that are impractical, but this is how progress happens.

    People will keep plugging away at this and figure out how to avoid wearing the hard drive, how to make it run faster, custom hardware buses etc.

    Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.

    1. arjie · · focus · HN ↗
      Haha 1T on $50k might be a bit hopeful, mate, even at FP8. But I too am hopeful.
      1. hedora · · focus · HN ↗
        AMD already demonstrated 1T on strix halo clusters. << $10K at original MSRP.
        1. arjie · · focus · HN ↗
          Problem is prefill on these, right? Initial prompt processing takes forever? I suppose you’re right. Cost is not a thing on its own. It’s a performance-cost frontier and one can do CPU inference in the worst case.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.