‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. IronWolve · · focus · HN ↗
    Getting about 200 tok/s on a 5090 with 64 gigs of ram, running the swift iq2_xs quant, ukisai/Swift-1.5-Qwen3.8-Flash-Next-GSQ-RCO-GGUF with strata. There are other smaller flash next quants if you have 8gig vram too.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.