‹ BackHN Continuity

Thread

Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

310 points · 140 comments · leonickson

  1. AHASIC · · focus · HN ↗
    I read a comment on here a few months back I wanna restate. Basically, there is a good chance that Apple is betting that the LLMs in the future will be so efficient that those that consumers will use everyday will be easily computed by the iPhone or even bigger ones on Macs. Honestly makes the most sense that we are heading that way in a few years latest.
    1. Mistletoe · · focus · HN ↗
      What hardware advances would we need to see for that to happen? It feels like everything in that arena has kind of plateaued.
      1. bobbylarrybobby · · focus · HN ↗
        The models themselves have far from plateaued. Maybe someone finds a way to get a really capable model down to, say, 12GB of ram. Then we'd be in business.
        1. swiftcoder · · focus · HN ↗
          Agreed. We've just seen DeepSeek post-train their ~300 billion parameter flash model to outperform their 1.6 trillion parameter pro model, in the space of a few months. There would seem to still be quite a few opportunities on the table to bring big model smarts down to the smaller models
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.