Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
snehesht · · focus · HN ↗
<a href="https://huggingface.co/Qwen/Qwen3.8-Flash-Next" rel="nofollow">https://huggingface.co/Qwen/Qwen3.8-Flash-Next
proc0 · · focus · HN ↗
incognito124 · · focus · HN ↗
snehesht · · focus · HN ↗
nicce · · focus · HN ↗
DoctorOetker · · focus · HN ↗
Bnjoroge · · focus · HN ↗
mickeyp · · focus · HN ↗
It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.
snehesht · · focus · HN ↗
JokerDan · · focus · HN ↗
gruturo · · focus · HN ↗
Also, quantization techniques have improved - the I in IQ3 stands for imatrix - Importance Matrix - it is a bit more surgical in what it cuts. The result is a model where the most important weights are even Q6 or above, the least important Q2 or even below, overall it takes the space of a Q3 but with better results.
Tade0 · · focus · HN ↗
latentsea · · focus · HN ↗
roscas · · focus · HN ↗
But this Qwen 3.8 Flash next coder is amazing running with Strata.
thatsabadlook · · focus · HN ↗
geye1234 · · focus · HN ↗
PcChip · · focus · HN ↗
What inference engine are you using for flash next?
anon373839 · · focus · HN ↗
Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)
geye1234 · · focus · HN ↗
It always detects its spelling mistakes, btw, but it worried me. It may turn 'rm -rf ' into 'rm -rf /' one day.
Almost certainly the problem is my config, not the image.
xiconfjs · · focus · HN ↗
thatsabadlook · · focus · HN ↗
a11r · · focus · HN ↗
swozey · · focus · HN ↗
I've been waiting for a 35b of 3.8, I don't really know what the other versions are about. I'm on 5g so juggling 40gb of model files sucks. And honestly I'm sick of tweaking this stuff for no, very little, or break-it level improvements. Qwen3.6-a35b has been solid for work, just don't give it freedom to wipe your data.
ENGNR · · focus · HN ↗
I’ve heard a quantised version of flash next can fit in ~50 gb of vram (which needs a system level flag set to go over 48gb)
But the m1 cpu is itself a bottleneck on prefill compared to say an m5, there’s no real getting around it. And the 400mb/s bandwidth starts to hurt without MOE
Hoping these model optimisations can see us through to 2028 because for everything other than LLMs this hardware is still over specced and working incredibly well