Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
snehesht · · focus · HN ↗
<a href="https://huggingface.co/Qwen/Qwen3.8-Flash-Next" rel="nofollow">https://huggingface.co/Qwen/Qwen3.8-Flash-Next
proc0 · · focus · HN ↗
incognito124 · · focus · HN ↗
snehesht · · focus · HN ↗
nicce · · focus · HN ↗
DoctorOetker · · focus · HN ↗
Bnjoroge · · focus · HN ↗
mickeyp · · focus · HN ↗
It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.
snehesht · · focus · HN ↗
JokerDan · · focus · HN ↗
gruturo · · focus · HN ↗
Also, quantization techniques have improved - the I in IQ3 stands for imatrix - Importance Matrix - it is a bit more surgical in what it cuts. The result is a model where the most important weights are even Q6 or above, the least important Q2 or even below, overall it takes the space of a Q3 but with better results.
ranguna · · focus · HN ↗
Tade0 · · focus · HN ↗
latentsea · · focus · HN ↗
roscas · · focus · HN ↗
But this Qwen 3.8 Flash next coder is amazing running with Strata.