‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

806 points · 356 comments · snehesht

  1. AntiRush · · focus · HN ↗
    I've been working on support for this model in ds4 on the RTX 6000 pro - it's been really great for my use cases. The ds4 q4 quant performs a lot better than other similar sizes that I've seen.

    Using the Q4 quant on an RTX 6000 Pro Workstation Edition at 450 watts:

      Code: prefill 1,251 tok/s decode 255.26 tok/s
      Prose: prefill 1,251 tok/s decode 198.78 tok/s 
    
    Most important for me, I can run 4 concurrent streams at 400+ tok/s.

    <a href="https:&#x2F;&#x2F;github.com&#x2F;fairfieldt&#x2F;ds4" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;fairfieldt&#x2F;ds4

    1. jacquesm · · focus · HN ↗
      DS4 is an odd model. I have it working on way too many GPUs and yet for many tasks Qwen 3.8 will do much better. It also tends to loop, which is super annoying.

      GLM5.3 runs on similar hardware and is much better so if you&#x27;re going to burn cycles and brain power on this maybe look at GLM5.3 as a comparison as well?

      Other than that, when you&#x27;re done with that card...

      1. segmondy · · focus · HN ↗
        ds4 is the inference engine not the model.
        1. jacquesm · · focus · HN ↗
          Sorry, but no, I use DS4 as the shorthand for Deepseek 4, it is no coincidence that that particular inference engine was called DS4.

          From the homepage of the inference engine:

          &quot;DwarfStar aims to be the best way to run a few excellent large language models on consumer hardware (that is, hardware that people can actually own). To reach this goal, we are building a small native inference engine optimized first for DeepSeek V4 Flash (including the experimental vision model), DeepSeek V4.1 Flash (Metal, and text inference on CUDA), and additionally GLM 5.2 and 5.3, GLM 5.3 Flash and DeepSeek V4 PRO, and Qwen3.8 Flash Next (Metal and CUDA).&quot;

          The model has been around longer than that. DS3 = Deepseek 3 etc. I think Salvatore named it pretty cleverly but he doesn&#x27;t automatically get to own a two letter acronym.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.