Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
Unofficial Hacker News client; not affiliated with Y Combinator.
jbird99 · · focus · HN ↗
wat10000 · · focus · HN ↗
Alpha3031 · · focus · HN ↗
wat10000 · · focus · HN ↗
sudo_cowsay · · focus · HN ↗
petu · · focus · HN ↗
Practically if you're not streaming weights 24/7 from a full SSD, then it shouldn't be a problem.
zozbot234 · · focus · HN ↗
petu · · focus · HN ↗
It claims that each individual page read induces read disturb across whole block. And references <a href="https://arxiv.org/pdf/2501.02517" rel="nofollow">https://arxiv.org/pdf/2501.02517 that tested Samsung 3D TLC and found ~518K sequential page reads in a block to be ECC threshold (although it's unclear how they got 518K number -- e.g. is it single worst chip they've tried? authors brings up 160 chip sample size later on).
With 7704 pages in a block that's only ~70 sequential block reads till data is lost and to retain data controller would have to refresh block fair bit earlier.. basically it gives modern 3D TLC SSD lifespan measured in months (1TB drive 24/7 sequential reads at 5GB/s).