Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
XCSme · · focus · HN ↗
On low, 3.8 Flash Next is better AND considerably faster.
My results[0][1], on a RTX 3090 + 128GB DDR4 RAM. I still have to see if I can optimize any more settings for either, but I think I will switch from 3.8 27b to 3.8 Flash Next for my local LLM uses.
[0]: <a href="https://aibenchy.com/?q=qwen+3.8+27b%2C+qwen+3.8+next" rel="nofollow">https://aibenchy.com/?q=qwen+3.8+27b%2C+qwen+3.8+next
[1]: <a href="https://aibenchy.com/compare/qwen-qwen3-8-27b-low/qwen-qwen3-8-flash-next-low/qwen-qwen3-8-27b-high/qwen-qwen3-8-flash-next-xhigh/" rel="nofollow">https://aibenchy.com/compare/qwen-qwen3-8-27b-low/qwen-qwen3...
kelvie · · focus · HN ↗
And that Qwen 27b outperforms both when set to high?
What quants are being compared here?
XCSme · · focus · HN ↗
The exact quant is mentioned on model page[0] IQ3_S, and I think 27b was Q4, via Ollama, the one that fits on a 3090 24GB
[0]: <a href="https://aibenchy.com/model/qwen-qwen3-8-flash-next-low/" rel="nofollow">https://aibenchy.com/model/qwen-qwen3-8-flash-next-low/