‹ BackHN Continuity

Thread

Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

310 points · 140 comments · leonickson

  1. AHASIC · · focus · HN ↗
    I read a comment on here a few months back I wanna restate. Basically, there is a good chance that Apple is betting that the LLMs in the future will be so efficient that those that consumers will use everyday will be easily computed by the iPhone or even bigger ones on Macs. Honestly makes the most sense that we are heading that way in a few years latest.
    1. Mistletoe · · focus · HN ↗
      What hardware advances would we need to see for that to happen? It feels like everything in that arena has kind of plateaued.
      1. harrouet · · focus · HN ↗
        I could definitely image Apple embedding a kind of LLM-optimized FPGA: slow to load (update) an LLM, but blazing fast at computing tokens.

        Who needs memory when your model is set in silicon ?

        1. KeplerBoy · · focus · HN ↗
          You don't an FPGA if you're taping out your own chips. But that is just a MMA accelerator with decent memory bandwidth. No secret sauce here.
          1. harrouet · · focus · HN ↗
            I am talking about reconfigurable gates to implement an LLM in silicon, i.e. an FPGA...
            1. xprnio · · focus · HN ↗
              What might the “parameters/layers to gates” ratio look like? My naive and uninformed guess would be that 1B+ parameter would also need a 1B+ gate FPGA, but according to google they typically range from tens of thousands to several million (which would still be a fraction of a billion).
        2. dgently7 · · focus · HN ↗
          why would apple make a chip that could be updated to improve the model when they could just sell you a better chip in the next years device?

          on device llm gives apple the new "better camera" "better screen" race they need to keep people coming back for the latest.

          for average users everything else is tapped out... screens, cameras wifi... all the core stuff is good enough now its hard to feel/see the difference model year to model year. embedded llm would let them ship something new and the on device ecosystem advantage is huge. especially as the gpt and claudes get ads and enshittified... the apple on device even if its less "capable" would be so compelling.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.