‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

805 points · 355 comments · snehesht

  1. a11r · · focus · HN ↗
    I&#x27;m a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I&#x27;m running 4-bit quants on an RTX Pro 6000 rented for approximately $1&#x2F;hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK
    1. Winfred-zz · · focus · HN ↗
      I just ran a set of benchmarks, ninfer-3090-qwen3.8-27b (so mix of Q4 and Q5) vs strata-qwen3.8-flash-next-iq3_xxs (so Q3):

      │------------------- │ Ninfer-3090 │ Strata

      │ Code generation │ 52&#x2F;78 (66.7%) │ 70&#x2F;78 (89.7%)

      │ Code completion │ 40&#x2F;50 (80.0%) │ 44&#x2F;50 (88.0%)

      │ Total------------- │ 92&#x2F;128 (71.9%) │ 114&#x2F;128 (89.1%)

      │ API failures------ │ 10 │ 5

      - Ninfer generation: ~122 min total.

      - Strata generation: ~142 min total.

      So strata is a little slower, but keep in mind that ninfer-3090 is very optimized for a Qwen 3.8. Standard Qwen 3.8 runs at 20 t&#x2F;s, this modified version can do 50 t&#x2F;s (but it&#x27;s extremely long in it&#x27;s thinking, it just goes on and on.

      This is on a 3090 that will crash unless power capped, with a Zen 2 CPU, 64GB DDR4 with a PCIe that refuses to go higher than 8x (basically pretty crappy all in all).

      Yet with some tweaking and optimizing I still manage to get strata to run at 40 to 60 t&#x2F;s.

      That strata has been optimized on my Oh My Pi conversations. So when I&#x27;m using it, it&#x27;s probably faster and closer to ninfer in speed than during those unoptimized benchmark tests.

      1. greenavocado · · focus · HN ↗
        In my private C coding benchmark, Strata&#x27;s Qwen Flash Next Coder IQ1_M beats Luna low and runs at 125 tok&#x2F;s on my 5090
        1. seviu · · focus · HN ↗
          Not as fast for me. I was able to snatch a spark before the horrible price hike, and I have it at 40&#x2F;50tok&#x2F;sec. Before I had the great 3090 setup which is what led me to get the spark. That 3090 is now relegated to creating videos of my kids doing dumb things.

          I now backordered two more sparks, before the hike... hoping they will honor the agreed price and dont cancel on me. They should arrive in a month,

          Its evidently clear the frontier labs won&#x27;t keep on giving us cheap inference for much longer. I also watch in disbelief how people say a sub is cheaper, which probably is. But you loose on so many other things (privacy, predictability, censorship...)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.