‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

819 points · 362 comments · snehesht

  1. snehesht · · focus · HN ↗
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next

    1. proc0 · · focus · HN ↗
      Do you know how it compares to Qwen 3.8 27B? I really want to compare the distilled ones with harness versus the full MoE versions.
      1. thatsabadlook · · focus · HN ↗
        Significantly better for both performance and real world use case. 3.8 27b is a good small model. This is a good model.
        1. geye1234 · · focus · HN ↗
          I find 27B more accurate -- maybe because I&#x27;m running at FP8 instead of NVFP4? Flash Next starts making spelling mistakes when I get to 150K context or so. Also it sometimes ignores .md file instructions. Not sure if others have found that.
          1. PcChip · · focus · HN ↗
            Spelling mistakes?

            What inference engine are you using for flash next?

            1. geye1234 · · focus · HN ↗
              I&#x27;m running Pennyroyal&#x27;s Docker image (on Podman) which uses sglang. I have a single RTX 6000 Blackwell and 128GB RAM. I turned off disk caching. I&#x27;m running with a ~500K context, but have been limiting it to 256K in the client (pi).

              It always detects its spelling mistakes, btw, but it worried me. It may turn &#x27;rm -rf &#x27; into &#x27;rm -rf &#x2F;&#x27; one day.

              Almost certainly the problem is my config, not the image.

              1. xiconfjs · · focus · HN ↗
                Did you check your GPU for memory errors&#x2F;defects?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.