‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. a11r · · focus · HN ↗
    I&#x27;m a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I&#x27;m running 4-bit quants on an RTX Pro 6000 rented for approximately $1&#x2F;hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK
    1. Winfred-zz · · focus · HN ↗
      I just ran a set of benchmarks, ninfer-3090-qwen3.8-27b (so mix of Q4 and Q5) vs strata-qwen3.8-flash-next-iq3_xxs (so Q3):

      │------------------- │ Ninfer-3090 │ Strata

      │ Code generation │ 52&#x2F;78 (66.7%) │ 70&#x2F;78 (89.7%)

      │ Code completion │ 40&#x2F;50 (80.0%) │ 44&#x2F;50 (88.0%)

      │ Total------------- │ 92&#x2F;128 (71.9%) │ 114&#x2F;128 (89.1%)

      │ API failures------ │ 10 │ 5

      - Ninfer generation: ~122 min total.

      - Strata generation: ~142 min total.

      So strata is a little slower, but keep in mind that ninfer-3090 is very optimized for a Qwen 3.8. Standard Qwen 3.8 runs at 20 t&#x2F;s, this modified version can do 50 t&#x2F;s (but it&#x27;s extremely long in it&#x27;s thinking, it just goes on and on.

      This is on a 3090 that will crash unless power capped, with a Zen 2 CPU, 64GB DDR4 with a PCIe that refuses to go higher than 8x (basically pretty crappy all in all).

      Yet with some tweaking and optimizing I still manage to get strata to run at 40 to 60 t&#x2F;s.

      That strata has been optimized on my Oh My Pi conversations. So when I&#x27;m using it, it&#x27;s probably faster and closer to ninfer in speed than during those unoptimized benchmark tests.

      1. airspresso · · focus · HN ↗
        &gt; it&#x27;s extremely long in it&#x27;s thinking, it just goes on and on

        Known issue with this model, I recommend setting thinking to &#x27;medium&#x27; instead of default &#x27;xhigh&#x27;.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.