‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. Tepix · · focus · HN ↗
    All headlines about LLM performance MUST have the quantization also mentioned in the headline.

    You know, so you're not wasting your time like in this post.

    1. latentsea · · focus · HN ↗
      This wasn't a waste of time for me at all. On llama.cpp I could only run the IQ3_XXS quant at 21 t/s on my R9700 + system RAM, and on Strata I can run it at 60 t/s. Also... QFN at IQ3_XXS is giving better results for me than 27B at Q6 fwiw.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.