‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. a11r · · focus · HN ↗
    I&#x27;m a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I&#x27;m running 4-bit quants on an RTX Pro 6000 rented for approximately $1&#x2F;hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK
    1. robot_jesus · · focus · HN ↗
      Can you say more about where you&#x27;re renting the RTX Pro 6000 for $1&#x2F;hour?
      1. a11r · · focus · HN ↗
        I am renting spot VMs from Nebius. I&#x27;ve tried a variety of other providers like Vast and Spheron. Vast worked well for renting 5090s but I like the large memory and pricing I&#x27;m getting at Nebius for RTX Pro 6000. The extra RAM really matters because I need PLE to offload the ngram to RAM.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.