‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

806 points · 356 comments · snehesht

  1. a11r · · focus · HN ↗
    I&#x27;m a little skeptical of going below 4-bit quants due to the potential for significant degradation in quality. I&#x27;m running 4-bit quants on an RTX Pro 6000 rented for approximately $1&#x2F;hour and getting about 1.2 million tokens out and 40 million tokens in per hour with caching. The quality of 4-bit quant is good enough for difficult but well-scoped coding tasks. Here is the inference stack I am using: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;BlackwellPerformance&#x2F;s&#x2F;FrKwk3GoDK
    1. Winfred-zz · · focus · HN ↗
      I just ran a set of benchmarks, ninfer-3090-qwen3.8-27b (so mix of Q4 and Q5) vs strata-qwen3.8-flash-next-iq3_xxs (so Q3):

      │------------------- │ Ninfer-3090 │ Strata

      │ Code generation │ 52&#x2F;78 (66.7%) │ 70&#x2F;78 (89.7%)

      │ Code completion │ 40&#x2F;50 (80.0%) │ 44&#x2F;50 (88.0%)

      │ Total------------- │ 92&#x2F;128 (71.9%) │ 114&#x2F;128 (89.1%)

      │ API failures------ │ 10 │ 5

      - Ninfer generation: ~122 min total.

      - Strata generation: ~142 min total.

      So strata is a little slower, but keep in mind that ninfer-3090 is very optimized for a Qwen 3.8. Standard Qwen 3.8 runs at 20 t&#x2F;s, this modified version can do 50 t&#x2F;s (but it&#x27;s extremely long in it&#x27;s thinking, it just goes on and on.

      This is on a 3090 that will crash unless power capped, with a Zen 2 CPU, 64GB DDR4 with a PCIe that refuses to go higher than 8x (basically pretty crappy all in all).

      Yet with some tweaking and optimizing I still manage to get strata to run at 40 to 60 t&#x2F;s.

      That strata has been optimized on my Oh My Pi conversations. So when I&#x27;m using it, it&#x27;s probably faster and closer to ninfer in speed than during those unoptimized benchmark tests.

      1. sleight42 · · focus · HN ↗
        Anecdotally, I&#x27;ve found that 27B at 4bit hallucinates a lot more than QFN at 3_xxs, particularly for less technical reasoning.

        I&#x27;ve tried to use both for online comparison shopping. QFN not only seemed less delusional but also made useful observations and problem solved ways around many different website access issues.

      2. tharkun__ · · focus · HN ↗
        You can use 4 spaces on HN to get code&#x2F;monospaced formatting:

            │------------------- │ Ninfer-3090    │ Strata    
            │ Code generation    │ 52&#x2F;78  (66.7%) │ 70&#x2F;78   (89.7%)
            │ Code completion    │ 40&#x2F;50  (80.0%) │ 44&#x2F;50   (88.0%)
            │ Total------------- │ 92&#x2F;128 (71.9%) │ 114&#x2F;128 (89.1%)
            │ API failures------ │ 10             │ 5
      3. greenavocado · · focus · HN ↗
        In my private C coding benchmark, Strata&#x27;s Qwen Flash Next Coder IQ1_M beats Luna low and runs at 125 tok&#x2F;s on my 5090
        1. seviu · · focus · HN ↗
          Not as fast for me. I was able to snatch a spark before the horrible price hike, and I have it at 40&#x2F;50tok&#x2F;sec. Before I had the great 3090 setup which is what led me to get the spark. That 3090 is now relegated to creating videos of my kids doing dumb things.

          I now backordered two more sparks, before the hike... hoping they will honor the agreed price and dont cancel on me. They should arrive in a month,

          Its evidently clear the frontier labs won&#x27;t keep on giving us cheap inference for much longer. I also watch in disbelief how people say a sub is cheaper, which probably is. But you loose on so many other things (privacy, predictability, censorship...)

      4. airspresso · · focus · HN ↗
        &gt; it&#x27;s extremely long in it&#x27;s thinking, it just goes on and on

        Known issue with this model, I recommend setting thinking to &#x27;medium&#x27; instead of default &#x27;xhigh&#x27;.

      5. seviu · · focus · HN ↗
        Unrelated... just use your just locally installed qwen 3.8 and instruct it to figure out why you arent at 16x.

        For me, I didnt know I was on 4x and I was able to, giving it enough permissions, with a good harness like PI, to figure things out and go 4x -&gt; 8x -&gt; 16x.

        It probably will take some opening the case and switching ssds around, but it&#x27;s amazing how powerful these models have become.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.