‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

793 points · 354 comments · snehesht

  1. snehesht · · focus · HN ↗
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next

    1. proc0 · · focus · HN ↗
      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
      1. incognito124 · · focus · HN ↗
        Qwen 3.8 flash next is way better than 27B. It&#x27;s so good I dont even use claude anymore
        1. snehesht · · focus · HN ↗
          Yeah I agree, I&#x27;m running it with Pi didn&#x27;t notice much difference compared to lower tier models and the speed, of course.
          1. nicce · · focus · HN ↗
            I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.
            1. DoctorOetker · · focus · HN ↗
              do LLMs tend to be homesick when not used in the same harness they sat in during some training phase?
              1. Bnjoroge · · focus · HN ↗
                iirc there was a sectionin Qwen’s paper where they talked anout how they post-trained flash or 3.8 to work just as well regardless of the harness or eval used. I think that used to be true but not sure if it is any longer
        2. mickeyp · · focus · HN ↗
          I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

          It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

          1. snehesht · · focus · HN ↗
            This is interesting, thanks. - <a href="https:&#x2F;&#x2F;github.com&#x2F;Neroued&#x2F;ninfer" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Neroued&#x2F;ninfer
        3. JokerDan · · focus · HN ↗
          Is this true for 27b Q4_K_XL vs flash next IQ3_S? I thought under Q4 models start quickly degrading?
          1. gruturo · · focus · HN ↗
            While this is generally true, it&#x27;s _a little_ less true the larger the model is.

            Also, quantization techniques have improved - the I in IQ3 stands for imatrix - Importance Matrix - it is a bit more surgical in what it cuts. The result is a model where the most important weights are even Q6 or above, the least important Q2 or even below, overall it takes the space of a Q3 but with better results.

            1. ranguna · · focus · HN ↗
              A lot of words, but no benchmarks. I&#x27;m a little tired of all the &quot;finger in the air&quot; vibe checks. You can say all you want, but you&#x27;ll only know once you put out some numbers. Which people have in other threads, and 4 bit 27B beats Flash at 3 bit
          2. Tade0 · · focus · HN ↗
            To add to the other comment, there&#x27;s also Ridge quantisation - the majority of weights are indeed Q3_x, but the most sensitive layers are FP8.
          3. latentsea · · focus · HN ↗
            This is the conventional wisdom, but in practice what matters is how reliably the model performs on your tasks in the real world. I have an R9700 and an RTX 5060 Ti and I&#x27;ve been running an IQ3_S quant of 27B on the 5060 Ti vs a Q6 quant on the R9700. I still manage to get stuff done with the IQ3_S quant.
        4. roscas · · focus · HN ↗
          I prefer <a href="https:&#x2F;&#x2F;ornith.ai&#x2F;ornith_1_5.html" rel="nofollow">https:&#x2F;&#x2F;ornith.ai&#x2F;ornith_1_5.html to Qwen 3.8 not only because it is much faster on my hardware but better responses.

          But this Qwen 3.8 Flash next coder is amazing running with Strata.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.