‹ BackHN Continuity

Thread

Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone

310 points · 140 comments · leonickson

  1. brrrrrm · · focus · HN ↗
    this is cool but like, are we just vibe coding NAND burners at this point? these decode times don't really tell the whole story, because prefill becomes the bottleneck.

    half an hour to process 10k tokens on an M5 seems... not great

    1. kennywinker · · focus · HN ↗
      Not great for coding, or realtime agent interactions. But for background processing tasks overnight? Seems like it’d work pretty well
      1. selcuka · · focus · HN ↗
        I'm pretty sure one can rent a GPU for a few minutes with the electricity cost of leaving an M5 overnight.
        1. hdgvhicv · · focus · HN ↗
          Domestic electricity is free nowadays, certainly for most of the year, as solar plus battery covers your usage for a tiny percentage of the cost of your house.
          1. wccrawford · · focus · HN ↗
            Only if you don't count the cost of the equipment and installation.
            1. hdgvhicv · · focus · HN ↗
              Or the cost of the house.

              Given the cost of a building is far more than the cost of generating enough power for that building it doesn’t really matter

        2. kennywinker · · focus · HN ↗
          Sure. One could. But then one wouldn’t be in control of every step of the process.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.