‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. snehesht · · focus · HN ↗
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next

    1. proc0 · · focus · HN ↗
      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
      1. incognito124 · · focus · HN ↗
        Qwen 3.8 flash next is way better than 27B. It&#x27;s so good I dont even use claude anymore
        1. snehesht · · focus · HN ↗
          Yeah I agree, I&#x27;m running it with Pi didn&#x27;t notice much difference compared to lower tier models and the speed, of course.
          1. nicce · · focus · HN ↗
            I am running 27B with Deepseek Harness these days and somehow just by using it, without any parameter changes, the model feels even more intelligent.
            1. DoctorOetker · · focus · HN ↗
              do LLMs tend to be homesick when not used in the same harness they sat in during some training phase?
              1. Bnjoroge · · focus · HN ↗
                iirc there was a sectionin Qwen’s paper where they talked anout how they post-trained flash or 3.8 to work just as well regardless of the harness or eval used. I think that used to be true but not sure if it is any longer
        2. mickeyp · · focus · HN ↗
          I have not tried Flash Next yet; but 27B is a cracking, little model. It is the first small model that I, as someone with 30 years of experience, can finally say is good enough to hand off small and mid-sized tasks and expect a pretty good result.

          It is also a competent tool caller when quantised to NVFP4 for use with ninfer; my own harness only reports the occasional hiccup and it is only because the model will sometimes emit tool calling tokens in its reasoning loop.

          1. snehesht · · focus · HN ↗
            This is interesting, thanks. - <a href="https:&#x2F;&#x2F;github.com&#x2F;Neroued&#x2F;ninfer" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;Neroued&#x2F;ninfer
        3. JokerDan · · focus · HN ↗
          Is this true for 27b Q4_K_XL vs flash next IQ3_S? I thought under Q4 models start quickly degrading?
          1. gruturo · · focus · HN ↗
            While this is generally true, it&#x27;s _a little_ less true the larger the model is.

            Also, quantization techniques have improved - the I in IQ3 stands for imatrix - Importance Matrix - it is a bit more surgical in what it cuts. The result is a model where the most important weights are even Q6 or above, the least important Q2 or even below, overall it takes the space of a Q3 but with better results.

          2. Tade0 · · focus · HN ↗
            To add to the other comment, there&#x27;s also Ridge quantisation - the majority of weights are indeed Q3_x, but the most sensitive layers are FP8.
          3. latentsea · · focus · HN ↗
            This is the conventional wisdom, but in practice what matters is how reliably the model performs on your tasks in the real world. I have an R9700 and an RTX 5060 Ti and I&#x27;ve been running an IQ3_S quant of 27B on the 5060 Ti vs a Q6 quant on the R9700. I still manage to get stuff done with the IQ3_S quant.
        4. roscas · · focus · HN ↗
          I prefer <a href="https:&#x2F;&#x2F;ornith.ai&#x2F;ornith_1_5.html" rel="nofollow">https:&#x2F;&#x2F;ornith.ai&#x2F;ornith_1_5.html to Qwen 3.8 not only because it is much faster on my hardware but better responses.

          But this Qwen 3.8 Flash next coder is amazing running with Strata.

      2. thatsabadlook · · focus · HN ↗
        Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model.
        1. geye1234 · · focus · HN ↗
          I find 27B more accurate -- maybe because I&#x27;m running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.
          1. PcChip · · focus · HN ↗
            Spelling mistakes?

            What inference engine are you using for flash next?

            1. anon373839 · · focus · HN ↗
              Yep, can confirm that is NOT normal. Are you using Nvidia’s NVFP4 quant? There are other NVFP4s floating around but they are not as good. The quality of the calibration data really matters.

              Qwen Flash Next is just excellent, all the way to the very end of the native 262k context. (I haven’t tried YaRN scaling to 1M, so I don’t know about that.)

            2. geye1234 · · focus · HN ↗
              I&#x27;m running Pennyroyal&#x27;s Docker image (on Podman) which uses sglang. I have a single RTX 6000 Blackwell and 128GB RAM. I turned off disk caching. I&#x27;m running with a ~500K context, but have been limiting it to 256K in the client (pi).

              It always detects its spelling mistakes, btw, but it worried me. It may turn &#x27;rm -rf &#x27; into &#x27;rm -rf &#x2F;&#x27; one day.

              Almost certainly the problem is my config, not the image.

              1. xiconfjs · · focus · HN ↗
                Did you check your GPU for memory errors&#x2F;defects?
          2. thatsabadlook · · focus · HN ↗
            Definitely not.
      3. a11r · · focus · HN ↗
        We recently moved from 27B to Flash Next. The quality is superior for coding. Our workload is primarily well-defined coding tasks that need to be attempted a few times before the model gets it just right. FlashNext is also better at finding issues in generated code than Gemini 3.8 Flash.
        1. swozey · · focus · HN ↗
          I&#x27;m on m1 max 64gb and went from qwen3.8-27B back to qwen3.6-a35b. Is flash next the move? I went from usable say 40tk&#x2F;s qwen3.6 to unusable, like 11 with 3.8 and not impressed with the replies for the time sacrifice. pi (omp) and omlx but not with the recent 3.8 patch.

          I&#x27;ve been waiting for a 35b of 3.8, I don&#x27;t really know what the other versions are about. I&#x27;m on 5g so juggling 40gb of model files sucks. And honestly I&#x27;m sick of tweaking this stuff for no, very little, or break-it level improvements. Qwen3.6-a35b has been solid for work, just don&#x27;t give it freedom to wipe your data.

          1. ENGNR · · focus · HN ↗
            Exact same scenario here

            I’ve heard a quantised version of flash next can fit in ~50 gb of vram (which needs a system level flag set to go over 48gb)

            But the m1 cpu is itself a bottleneck on prefill compared to say an m5, there’s no real getting around it. And the 400mb&#x2F;s bandwidth starts to hurt without MOE

            Hoping these model optimisations can see us through to 2028 because for everything other than LLMs this hardware is still over specced and working incredibly well

    2. thatsabadlook · · focus · HN ↗
      Why is this surprisingly well? It&#x27;s 2.5x faster than anthropic models, you have data sovereignty, privacy,and that&#x27;s a strong model. Sounds like a best case scenario to me
      1. hdjrudni · · focus · HN ↗
        Not sure you understand the term &#x27;surprisingly well&#x27;. It means &#x27;better than expected&#x27;. I suspect they parent poster didn&#x27;t actually expect to get &gt;= 100 T&#x2F;s.
    3. roscas · · focus · HN ↗
      Coder version with 30t&#x2F;sec on a Ryzen 3600x with 48GB of RAM with a nvidia 3080.

      This is not a very fast desktop. Memory speed is around 2000mhz only. My SSD is some of the worst SSD I&#x27;ve seen and 3080 had its days of glory.

      I still have code, chromium, librewolf and many other programs running. I have video streams running while I also watch tv and many times youtube videos.

      I use it with the browser that has a great dashboard and with hermes agent and that it really makes this amazing.Only change I made is to set thinking to low.

      This is a coding model. Any other task, I still use Ornith 1.5 35B that throws 20t&#x2F;sec and Laguna.XS-2.0.

      1. StumpChunkman · · focus · HN ↗
        How much VRAM on your 3080? I&#x27;ve got an early 10gb model. I&#x27;ve been thinking of exploring local coding models, but everyone seems to use much better GPUs than I have access to. Yours is one of the first I&#x27;ve seen with maybe similar hardware on some level.
        1. roscas · · focus · HN ↗
          Yes, 3080 with 10GB, forgot to mention that.

          Mine is at the moment writting some cpp code for some SBOM tests.

          I have loads of terminals open. Librewolf, Chromium and you know how this crap likes ram, I have also a vm with 4gb of ram running and doing stuff while I wait for the results but hey, while I wrote this the program is done. Wow! That was 29.x tokens per second most of the time.

          Oh I will run some other tests with hermes now because hermes is amazing too.

        2. Abishek_Muthian · · focus · HN ↗
          With limited VRAM, you can still use local LLMs but you have to set your expectations correctly to solve the right problem.

          I run small models on various kind of devices including 1.7B model on the original Jetson Nano (4B) abandoned by Nvidia, I had to upscale the software (OS, Lllama.cpp etc.) to run the model but the model runs at 17t&#x2F;s.

          Small models are great at NLP stuff like classification (e.g. bookmarking), ASR etc. I have built custom browser extensions to save time with bookmarking and categorizing to use with local LLMs.

      2. la_oveja · · focus · HN ↗
        so it runs on 10gb vram?
    4. notnullorvoid · · focus · HN ↗
      Which quantization are you using to reach those numbers?
    5. jacquesm · · focus · HN ↗
      Speed is one thing, accuracy another. Have you benchmarked it against a reference? If so, what were the results? I tend to go for accuracy over speed because usually that means fewer round trips and fewer tokens wasted.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.