‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. AntiRush · · focus · HN ↗
    I've been working on support for this model in ds4 on the RTX 6000 pro - it's been really great for my use cases. The ds4 q4 quant performs a lot better than other similar sizes that I've seen.

    Using the Q4 quant on an RTX 6000 Pro Workstation Edition at 450 watts:

      Code: prefill 1,251 tok/s decode 255.26 tok/s
      Prose: prefill 1,251 tok/s decode 198.78 tok/s 
    
    Most important for me, I can run 4 concurrent streams at 400+ tok/s.

    <a href="https:&#x2F;&#x2F;github.com&#x2F;fairfieldt&#x2F;ds4" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;fairfieldt&#x2F;ds4

    1. jacquesm · · focus · HN ↗
      DS4 is an odd model. I have it working on way too many GPUs and yet for many tasks Qwen 3.8 will do much better. It also tends to loop, which is super annoying.

      GLM5.3 runs on similar hardware and is much better so if you&#x27;re going to burn cycles and brain power on this maybe look at GLM5.3 as a comparison as well?

      Other than that, when you&#x27;re done with that card...

      1. AntiRush · · focus · HN ↗
        Agreed that glm-5.3-flash is a strong model. On a single 6000 pro you can (barely) fit a q2 quant in vram. I haven&#x27;t used it extensively, but my initial feeling is that the q2 is quite a bit worse than q4 for this model. To run it comfortably at q4 with reasonable context length you really need 2 6000s.

        When the qwen 4 series is released I am hopeful there&#x27;ll be a strong model with the same architecutre. as qwen3.8-flash-next.

        1. jacquesm · · focus · HN ↗

            |  0  N&#x2F;A  N&#x2F;A  321801 C llama-server 16540MiB 
            |  1  N&#x2F;A  N&#x2F;A  321801 C llama-server 22404MiB 
            |  2  N&#x2F;A  N&#x2F;A  321801 C llama-server 23480MiB 
            |  4  N&#x2F;A  N&#x2F;A  321801 C llama-server 43982MiB 
            |  5  N&#x2F;A  N&#x2F;A  321801 C llama-server 41656MiB 
            |  6  N&#x2F;A  N&#x2F;A  321801 C llama-server 22366MiB 
            |  7  N&#x2F;A  N&#x2F;A  321801 C llama-server 16208MiB
          
          That&#x27;s 3.86 bits per word I could run the 4 bpw one as well (there is still one spare gpu and another 27G on the ones listed above. DS4 uses about the same memory, is a little bit faster (though I suspect that by the time GLM 5.3 support is a bit more mature the speed difference will have evaporated).
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.