Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
esafak · · focus · HN ↗
I think publishing benchmarks with quantized models should become standard practice.
mkl · · focus · HN ↗
> Coder: a coding version with half of the experts removed. It reaches 91% of the full model's SWE-bench Verified score (measured by its authors) and fits 32 GB of RAM.
<a href="https://github.com/Niko1221/Strata#which-model-should-i-pick" rel="nofollow">https://github.com/Niko1221/Strata#which-model-should-i-pick
nicce · · focus · HN ↗
XCSme · · focus · HN ↗
Both 27b and flash next are more stable on "low" reasoning, only for generative /creative tasks, xhigh could be better, but both suffer from way too much reasoning at xhigh. And neither really support high, so low is the best reasoning effort.
[0]: <a href="https://aibenchy.com/compare/qwen-qwen3-8-27b-low/qwen-qwen3-8-flash-next-low/" rel="nofollow">https://aibenchy.com/compare/qwen-qwen3-8-27b-low/qwen-qwen3...