Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Show HN: Run an 80B Qwen in 4.3 GB of RAM on a Mac, and a 35B on an iPhone
Unofficial Hacker News client; not affiliated with Y Combinator.
dghlsakjg · · focus · HN ↗
People will keep plugging away at this and figure out how to avoid wearing the hard drive, how to make it run faster, custom hardware buses etc.
Keep going! I personally can't wait for the day when a 1t param model runs off a $200 SSD instead of a $50k rack of Nvidia chips.
arjie · · focus · HN ↗
hedora · · focus · HN ↗
arjie · · focus · HN ↗
apimade · · focus · HN ↗
1080 GTX in 2016. Cutting-edge, an insanely powerful consumer card for the time. Theoretically around 8.87 to 8.9 TFLOPS.
5090 RTX in 2026. Cutting-edge, an insanely powerful consumer card for today. Theoretically around 104.8 TFLOPS.
In the same timeframe mobile processor CPU's went from 0.001 TFLOPS, to today's Apple's A19 Pro chip which delivers 2.074 TFLOPS.
That's _without_ getting into ASIC's, or purpose-built hardware like Taalas's model on silicon HC1, or generic AI dies like what they're planning with HC2 or Cerebras, which will massively compress the timeline.
cududa · · focus · HN ↗
xbmcuser · · focus · HN ↗
formerly_proven · · focus · HN ↗
zmmmmm · · focus · HN ↗
jack_pp · · focus · HN ↗
Aren't we already approaching theoretical physical limits? We're at 2nm
kaashif · · focus · HN ↗
(2) Are you saying that you think we're at the limits of computing in general, or that specific technology?
We know, for example, that a human brain level intelligence is possible to run on a human brain. We are nowhere near that. And actually that's not even a physical limit necessarily.
But that is...not a low hanging fruit.
fragmede · · focus · HN ↗
Nowhere?
darkwater · · focus · HN ↗
cvak · · focus · HN ↗
root_axis · · focus · HN ↗
apimade · · focus · HN ↗
GTX 1080 in 2016: 8 GB of GDDR5X, with 320 GB/s.
RTX 5090 in 2026: 32 GB of GDDR7, with 1.792 TB/s.
This is fun, what's next?!
PCI 8.0 is breaking 1TB/s, GDDR7 is 1TB/s.
With just the _current_ timeline, things are looking like they'll compress once we get over this initial lump.
flaburgan · · focus · HN ↗
dtj1123 · · focus · HN ↗
The suggestion is that a 1T model could be made to run on cheap consumer hardware of the future.
foxrider · · focus · HN ↗
SJC_Hacker · · focus · HN ↗
At the rate models are improving, it would be obsolete in six months.
HPsquared · · focus · HN ↗
naasking · · focus · HN ↗
foxrider · · focus · HN ↗
kaelwd · · focus · HN ↗
apimade · · focus · HN ↗
This is definitely being done with private models by HFT/quant firms, data processing agencies/orgs (large intelligence agencies, _every_ data analytics org, etc).
Azantys · · focus · HN ↗
dghlsakjg · · focus · HN ↗
We went from adding 8 teraflops in a decade, to adding almost 100 the next decade. If we add "only" 400 more teraflops in the next decade the graph will make that initial growth look flat in comparison, even though your math would show that we are basically stalled out.
It’s like claiming that a company that goes from making $1 to $1k to $100k to $1mm in a 4 year period has decelerating growth.
ksec · · focus · HN ↗
Because it is decelerating growth. There is a reason why we use YoY percentage in annual and financial reporting.
sh3rl0ck · · focus · HN ↗
Not too wild an idea!
arjie · · focus · HN ↗