‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. Tepix · · focus · HN ↗
    Q2 quantization. Not interested.
    1. latentsea · · focus · HN ↗
      You can run IQ3_XXS, IQ3_S, and IQ4_XS too. I've switched to IQ3_XXS and am running at 60 t/s on Strata vs the 21 t/s I was getting in llama.cpp. Better outputs too.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.