‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. snehesht · · focus · HN ↗
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next

    1. proc0 · · focus · HN ↗
      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
      1. thatsabadlook · · focus · HN ↗
        Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model.
        1. geye1234 · · focus · HN ↗
          I find 27B more accurate -- maybe because I&#x27;m running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.
          1. PcChip · · focus · HN ↗
            Spelling mistakes?

            What inference engine are you using for flash next?

            1. anon373839 · · focus · HN ↗
              Yep, can confirm that is NOT normal. Are you using Nvidia’s NVFP4 quant? There are other NVFP4s floating around but they are not as good. The quality of the calibration data really matters.

              Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.