‹ BackHN Continuity

Thread

Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

310 points · 140 comments · leonickson

  1. dghlsakjg · · focus · HN ↗
    I know everyone wants to crap all over these setups that are impractical, but this is how progress happens.

    People will keep plugging away at this and figure out how to avoid wearing the hard drive, how to make it run faster, custom hardware buses etc.

    Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.

    1. marci · · focus · HN ↗
      Seems like what Apple's going for with afm3. Their latest model that will be embedded in macOS 27 is a quantized dense 20B that only select between 1 to 4B at inference, based on the prompt, not token by token. If only they could make a 100B or 400B dense that selects ~5 to 15B...
      1. josu · · focus · HN ↗
        I don't understand, if they are only using a subset of the tokens then it's a sparse model. What do you mean by dense?
        1. marci · · focus · HN ↗
          Nothing to understand. Straight up hallucination. I could have sworn I read that they used a novel architecture where the model is dense but you could select specific layers or something at inference. reread the announcement: just said MoE. Corrected my brain's weights so thanks.

          <a href="https:&#x2F;&#x2F;machinelearning.apple.com&#x2F;research&#x2F;introducing-third-generation-of-apple-foundation-models" rel="nofollow">https:&#x2F;&#x2F;machinelearning.apple.com&#x2F;research&#x2F;introducing-third...

          1. josu · · focus · HN ↗
            Thanks for responding.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.