‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

826 points · 365 comments · snehesht

  1. a11r · · focus · HN ↗
    I&#x27;m a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I&#x27;m running 4-bit quants on an RTX Pro 6000 rented for approximately $1&#x2F;hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK
    1. sudo_cowsay · · focus · HN ↗
      For the 1 dollar per hour, may I ask (without trying to incite things) why you would not just buy a ChatGPT Plus subscription or other 20 dollar subscriptions? Do you like the privacy?
      1. manmal · · focus · HN ↗
        Those plans are super slow right now. Nowhere near 100t&#x2F;s.
        1. martianvoid · · focus · HN ↗
          I think Opus 5.5 is now fast enough with the inference stack improvements that they did. Sonnet 5.5 is around 140 tokens&#x2F;s when I had measured it last time. A $20 subs on claude and using only sonnet 5.5 will last a lot
          1. c16 · · focus · HN ↗
            I was firmly wanting to move towards local llm, and since the Sonnet 5.5 and recent Opus speed increases, I&#x27;m putting that on hold until local inference speeds can be improved. Not sure how much juice there is to squeeze there, but I&#x27;m hopeful.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.