Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
esafak · · focus · HN ↗
I think publishing benchmarks with quantized models should become standard practice.
mkl · · focus · HN ↗
> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.
<a href="https://github.com/Niko1221/Strata#which-model-should-i-pick" rel="nofollow">https://github.com/Niko1221/Strata#which-model-should-i-pick
nicce · · focus · HN ↗
kennywinker · · focus · HN ↗
27b at q4 is ~16gb
So from a raw amount of data, qwen3.8-flash-next wins easily. But flash-next is an MoE model, so it only has 6b parameters active per token, vs 27b's dense 27b per token. So 27b@q4 uses ~16gb of weights per token, and flash-next uses about 4gb of weights (125/80 * 6).
But those numbers don't really tell us anything useful, because there is an interplay between total model size and active parameters and intelligence that isn't obvious or simple.
(sizes are based on the unsloth quants, not the coder variant, but the idea holds - this isn't calculatable with simple math, you gotta test them and see)
XCSme · · focus · HN ↗
Both 27b and flash next are more stable on "low" reasoning, only for generative /creative tasks, xhigh could be better, but both suffer from way too much reasoning at xhigh. And neither really support high, so low is the best reasoning effort.
[0]: <a href="https://aibenchy.com/compare/qwen-qwen3-8-27b-low/qwen-qwen3-8-flash-next-low/" rel="nofollow">https://aibenchy.com/compare/qwen-qwen3-8-27b-low/qwen-qwen3...