‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. snehesht · · focus · HN ↗
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next

    1. proc0 · · focus · HN ↗
      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
      1. incognito124 · · focus · HN ↗
        Qwen 3.8 flash next is way better than 27B. It&#x27;s so good I dont even use claude anymore
        1. JokerDan · · focus · HN ↗
          Is this true for 27b Q4_K_XL vs flash next IQ3_S? I thought under Q4 models start quickly degrading?
          1. latentsea · · focus · HN ↗
            This is the conventional wisdom, but in practice what matters is how reliably the model performs on your tasks in the real world. I have an R9700 and an RTX 5060 Ti and I&#x27;ve been running an IQ3_S quant of 27B on the 5060 Ti vs a Q6 quant on the R9700. I still manage to get stuff done with the IQ3_S quant.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.