‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. zkmon · · focus · HN ↗
    I don't get it. It's file size is about 6 times larger than 27B model for the same quant, but the performance improvement is hardly 10% across all benchmarks, according the metrics on it's hf page. Why should one devote so much more hardware for so little benefit?
    1. petu · · focus · HN ↗
      6B activated weights per token vs 27B. Something like DGX Spark is way better suited for Flash Next.
      1. anon373839 · · focus · HN ↗
        It is on paper, but crazy enough, both models at NVFP4 run similar speeds for decode! The reason is that much more sophisticated speculative drafting is available for 27B. I’m hoping this will come to Flash Next, but I know MoEs pose challenges with that.
    2. latentsea · · focus · HN ↗
      Maybe use it and find out. Previously I was only able to run Qwen3.8-27B at acceptable (to me) speeds of 35 t/s on my R9700, but with Strata I'm doing 60 t/s running Qwen3.8-Flash-Next IQ3_XXS. I'm getting better results...
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.