Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
Luker88 · · focus · HN ↗
Surprisingly useful as long as you can leave it running a couple of hours at the very least.
While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.
londons_explore · · focus · HN ↗
Everything else (weights) are shared amongst tens of thousands of users currently doing inference in that cluster, so even if there are terabytes of weights for the model, they aren't much on a per-user basis.
eurekin · · focus · HN ↗
Where's that figure comning from? Last time I checked (could be the 3.6 Qwen 27b) single token needed 32kb
jburgess777 · · focus · HN ↗
bitexploder · · focus · HN ↗
Luker88 · · focus · HN ↗