‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. a11r · · focus · HN ↗
    I&#x27;m a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I&#x27;m running 4-bit quants on an RTX Pro 6000 rented for approximately $1&#x2F;hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK
    1. NinjaTrance · · focus · HN ↗
      Just as curiosity, how long does it take to set the environment up and running?

      Is it viable to start&#x2F;stop it multiple times per day?

      1. a11r · · focus · HN ↗
        Yes, you can probably get the whole thing up and running in about an hour the first time. If you pause and restart, it takes about 15 minutes to load the models from disk into GPU memory, so budget for cold startup time.
        1. KeplerBoy · · focus · HN ↗
          How does it take fifteen minutes to read &lt;100 GB into GPU memory? Shouldn&#x27;t that be limited by SSD speed with everything slower than a minute being a terrible ssd?
          1. teaearlgraycold · · focus · HN ↗
            A lot of cloud platforms have terrible slow network storage. They also might need to compile the GPU kernels fresh as they might not have a persistent CUDA cache.
          2. boredatoms · · focus · HN ↗
            It also depends on the runtime, vllm is unbelievably slow at model loading compared to llama.cpp
            1. ycui7 · · focus · HN ↗
              this is not right. vllm can easily load safetensors model at &gt;5GB&#x2F;s if not faster when setup right. did you make the compile cache persist? if you use docker, you should bind mount the kernel compile cache, so they don&#x27;t need to be recompiled each time vllm restart.
      2. prettyblocks · · focus · HN ↗
        I just had claude code set it up for me and create some startup scripts. The longest part was the model download.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.