‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

793 points · 354 comments · snehesht

  1. snehesht · · focus · HN ↗
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next

    1. proc0 · · focus · HN ↗
      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
      1. a11r · · focus · HN ↗
        We recently moved from 27B to Flash Next. The quality is superior for coding. Our workload is primarily well-defined coding tasks that need to be attempted a few times before the model gets it just right. FlashNext is also better at finding issues in generated code than Gemini 3.8 Flash.
        1. swozey · · focus · HN ↗
          I&#x27;m on m1 max 64gb and went from qwen3.8-27B back to qwen3.6-a35b. Is flash next the move? I went from usable say 40tk&#x2F;s qwen3.6 to unusable, like 11 with 3.8 and not impressed with the replies for the time sacrifice. pi (omp) and omlx but not with the recent 3.8 patch.

          I&#x27;ve been waiting for a 35b of 3.8, I don&#x27;t really know what the other versions are about. I&#x27;m on 5g so juggling 40gb of model files sucks. And honestly I&#x27;m sick of tweaking this stuff for no, very little, or break-it level improvements. Qwen3.6-a35b has been solid for work, just don&#x27;t give it freedom to wipe your data.

          1. ENGNR · · focus · HN ↗
            Exact same scenario here

            I’ve heard a quantised version of flash next can fit in ~50 gb of vram (which needs a system level flag set to go over 48gb)

            But the m1 cpu is itself a bottleneck on prefill compared to say an m5, there’s no real getting around it. And the 400mb&#x2F;s bandwidth starts to hurt without MOE

            Hoping these model optimisations can see us through to 2028 because for everything other than LLMs this hardware is still over specced and working incredibly well

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.