‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. a11r · · focus · HN ↗
    I&#x27;m a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I&#x27;m running 4-bit quants on an RTX Pro 6000 rented for approximately $1&#x2F;hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK
    1. nialv7 · · focus · HN ↗
      ISTA-DASLab&#x27;s IQ3_S quant is really good <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;ISTA-DASLab&#x2F;Qwen3.8-Flash-Next-GSQ-RCO-GGUF" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;ISTA-DASLab&#x2F;Qwen3.8-Flash-Next-GSQ-RC...
      1. ricardobeat · · focus · HN ↗
        What are you running it on? I&#x27;m getting mixed results using the IQ3_XXS quant which supposedly matches baseline, it feels significantly degraded.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.