Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s
Unofficial Hacker News client; not affiliated with Y Combinator.
AntiRush · · focus · HN ↗
Using the Q4 quant on an RTX 6000 Pro Workstation Edition at 450 watts:
Most important for me, I can run 4 concurrent streams at 400+ tok/s.<a href="https://github.com/fairfieldt/ds4" rel="nofollow">https://github.com/fairfieldt/ds4
jacquesm · · focus · HN ↗
GLM5.3 runs on similar hardware and is much better so if you're going to burn cycles and brain power on this maybe look at GLM5.3 as a comparison as well?
Other than that, when you're done with that card...
AntiRush · · focus · HN ↗
When the qwen 4 series is released I am hopeful there'll be a strong model with the same architecutre. as qwen3.8-flash-next.
jacquesm · · focus · HN ↗