‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. mmaunder · · focus · HN ↗
    More great work on local model but you’re still losing a lot. Down to 2 bit quantization and the coder model throws away half the MoE experts. In a world where anything is better than nothing, this is a net win. But we have a way to go still.
    1. latentsea · · focus · HN ↗
      You can run IQ3_XXS, IQ3_S and the IQ4_XS quants on this too. It works. It's fantastic. I'm getting better results than 27B now.
      1. bitexploder · · focus · HN ↗
        To add: <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;ISTA-DASLab&#x2F;Qwen3.8-Flash-Next-GSQ-RCO-GGUF" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;ISTA-DASLab&#x2F;Qwen3.8-Flash-Next-GSQ-RC... IQ3_XXS is within a point of the fully unquantized model and IQ3_S actually beats the unquantized model on many tasks! You lose absolutely nothing. It is quantization magic :)
        1. Muromec · · focus · HN ↗
          No magic needed, the quantization found the expert which deals in java and deleted it, so the model overall became better.
          1. bitexploder · · focus · HN ↗
            lol. JDQ -- Java Delete Quantization it&#x27;s the new thing!
          2. latentsea · · focus · HN ↗
            Ah, I see you&#x27;ve found the AbstractExpertRemovalFactoryFactory!
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.