‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. snehesht · · focus · HN ↗
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next

    1. proc0 · · focus · HN ↗
      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
      1. incognito124 · · focus · HN ↗
        Qwen 3.8 flash next is way better than 27B. It&#x27;s so good I dont even use claude anymore
        1. roscas · · focus · HN ↗
          I prefer <a href="https:&#x2F;&#x2F;ornith.ai&#x2F;ornith_1_5.html" rel="nofollow">https:&#x2F;&#x2F;ornith.ai&#x2F;ornith_1_5.html to Qwen 3.8 not only because it is much faster on my hardware but better responses.

          But this Qwen 3.8 Flash next coder is amazing running with Strata.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.