‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. Luker88 · · focus · HN ↗
    Currently I am running llama-cpp with `Qwen3.8-Flash-Next-UD-IQ3_XXS` on an old ryzen 8845HS with 96G of ram (and no dedicated graphics card) at 7tk/s and ~60tk/s filling, max ~120K context window.

    Surprisingly useful as long as you can leave it running a couple of hours at the very least.

    While huge models will still be better I think the general availability of RAM might be the downfall of AI companies.

    1. bitexploder · · focus · HN ↗
      Highly recommend the RCO-GSQ quant by ITSA btw. At IQ3_XSS it is within one point of the fully unquantized model.
      1. Luker88 · · focus · HN ↗
        will try, thank you for the pointer!
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.