‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. AntiRush · · focus · HN ↗
    I've been working on support for this model in ds4 on the RTX 6000 pro - it's been really great for my use cases. The ds4 q4 quant performs a lot better than other similar sizes that I've seen.

    Using the Q4 quant on an RTX 6000 Pro Workstation Edition at 450 watts:

      Code: prefill 1,251 tok/s decode 255.26 tok/s
      Prose: prefill 1,251 tok/s decode 198.78 tok/s 
    
    Most important for me, I can run 4 concurrent streams at 400+ tok/s.

    <a href="https:&#x2F;&#x2F;github.com&#x2F;fairfieldt&#x2F;ds4" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;fairfieldt&#x2F;ds4

    1. anon373839 · · focus · HN ↗
      Those prefill numbers don’t seem right for Qwen Flash Next. I get 3,000 tok&#x2F;sec on a DGX Spark and that’s got a lot less raw compute than the 6000 Pro.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.