‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

806 points · 356 comments · snehesht

  1. mark_l_watson · · focus · HN ↗
    Qwen 3.8 Flash Next is amazing. I only have a 64G Mac so I have to run Sushi project’s 3 bit quant. Amazing results with pi-dev. More for fun than anything else, but I am trying to do as much as possible with local models, now rarely falling back to a paid deepseek-4.1-flash API.

    Progress on running local models has been amazing.

    1. generalizations · · focus · HN ↗
      I haven't seen much in the way of benchmarks of those smaller quants. How does it compare to e.g. various generations of Opus?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.