‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

793 points · 354 comments · snehesht

  1. 0xbadcafebee · · focus · HN ↗
    Lol, sure, if you quant it to hell (Q2) it'll go real fast...

    They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.

    It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.

    1. sigbottle · · focus · HN ↗
      It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference?
      1. nottorp · · focus · HN ↗
        Is Qwen 3.8 at Q4 good enough?

        I tried to run 3.5 27b Q4 on what local hardware i had (only 8 Gb) and i was very disappointed. 3.8 wouldn't have fit in my VRAM and i wasn't in the mood to leave it overnight at slow speeds so I didn't try.

        1. cycomanic · · focus · HN ↗
          I've run 3.8 flash next k4_xl on my Strix halo box (128GB). And in the work I have done so far it was not significantly worse than recent GPT (running default model on pro plan). Admittedly I was not doing complex work (reorganizing a jupyterbook), but I could not see significant difference in the quality of the work. It was a striking difference to Laguna s 2.1 which I had tried just before (much faster and much better quality).
          1. nottorp · · focus · HN ↗
            At Q what? That's what I'm mostly asking about.

            My little test was "generate me a single page tic tac toe game in plain javascript. computer always play O. add unbeatable minmax. have the board, a status line and a new game button'. I used both lm studio and whatever the name of their new coding assistant that supercedes lm studio is.

            Qwen 3.5 Q4 went into some kind of loop where it fixed whatever was broken on the previous iteration only to have it broken some other way. (I was writing the description of the errors).

            Paid $20/mo claude opus did it right the first time. Or at worst it fixed the code based on descriptions without entering a breakage loop, iForgot. I know it isn't fair because it has 1 million tokens but still, it was just tic tac toe.

            But since everyone says qwen is decent, it's either:

            - Q4 is too little

            - my idea of "decent" is too much

            - 3.8 is much better than 3.5 even at Q4

            1. Balinares · · focus · HN ↗
              3.8 is absolutely quite a bit better than 3.5. My go-to quick benchmark is a bit like yours, implementing a simple but complete game in pygame. Most small models output code that's too buggy to be fixable. Qwen 3.8 wrote buggy code too, but the bugs were minor, and the game legitimately playable and fun.

              That said, I've not found it very reliable, especially at the low quant required to run on my hardware. It's very capable for its size, enough to be legitimately useful, but it's never certain whether it will hit its capability peak on a given attempt, and sometimes you have to give it multiple tries.

        2. Zambyte · · focus · HN ↗
          3.8 is a huge step up from 3.5, quantized or not. I do all of my programming on Qwen 3.8 27B Q4 these days.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.