‹ BackHN Continuity

Thread

Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

310 points · 140 comments · leonickson

  1. dghlsakjg · · focus · HN ↗
    I know everyone wants to crap all over these setups that are impractical, but this is how progress happens.

    People will keep plugging away at this and figure out how to avoid wearing the hard drive, how to make it run faster, custom hardware buses etc.

    Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.

    1. arjie · · focus · HN ↗
      Haha 1T on $50k might be a bit hopeful, mate, even at FP8. But I too am hopeful.
      1. apimade · · focus · HN ↗
        8800 GTX in 2006. Cutting-edge, an insanely powered consumer card for the time. Theoretically around 0.3456 TFLOPS.

        1080 GTX in 2016. Cutting-edge, an insanely powerful consumer card for the time. Theoretically around 8.87 to 8.9 TFLOPS.

        5090 RTX in 2026. Cutting-edge, an insanely powerful consumer card for today. Theoretically around 104.8 TFLOPS.

        In the same timeframe mobile processor CPU's went from 0.001 TFLOPS, to today's Apple's A19 Pro chip which delivers 2.074 TFLOPS.

        That's _without_ getting into ASIC's, or purpose-built hardware like Taalas's model on silicon HC1, or generic AI dies like what they're planning with HC2 or Cerebras, which will massively compress the timeline.

        1. foxrider · · focus · HN ↗
          Speaking of ASICs - how likely is it that as models get better we'll see someone baking a whole model directly into the silicon? It's like having l0 cache.
          1. SJC_Hacker · · focus · HN ↗
            You could do it but there would be no point, The only advantage over would be power consumption. And it would be quite expensive.

            At the rate models are improving, it would be obsolete in six months.

            1. HPsquared · · focus · HN ↗
              Power consumption and latency are very important on mobile
              1. naasking · · focus · HN ↗
                They're important everywhere of course, but especially on mobile. If AI researchers figure out how to offload knowledge and expertise from reasoning weights, then a core reasoning ASIC linked to the knowledge would totally rock.
            2. foxrider · · focus · HN ↗
              Yes, right now it would be obsolete in six months, but I also must add that this never stopped crypto miners from making new ASICs. However, with how useful Kimi is right now - at some point if someone makes a dedicated hardware board with "good enough" model for daily tasks - that would be a very sought after commodity.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.