‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. 0xbadcafebee · · focus · HN ↗
    Lol, sure, if you quant it to hell (Q2) it'll go real fast...

    They even link to a Q1 quant (Qwen3.8-Flash-Next-GSQ-RCO-Coder-GGUF) with half the experts ripped out. The idea is it'll go much faster and supposedly benches to not-terrible results. But the problem is you can't rely on it for real world long-horizon coding because that's where reasoning comes in, which is why you want the other layers.

    It turns out there's still no free lunch. Either get enough VRAM for a Q4, or use a much smaller model. Lobotomizing a larger model just to say you can run it fast isn't useful.

    1. sigbottle · · focus · HN ↗
      It's interesting though that Q4 seems to be enough, is there a reason that 4 bit floats are good enough for inference?
      1. amelius · · focus · HN ↗
        3 is the magic number, and 4 > 3.

        (seriously, nobody knows why any of this works; it's just a matter of trying)

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.