‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. lsb · · focus · HN ↗
    There’s other slop projects to run of Qwen, like ds4, would be interesting to see a comparison
    1. happycube · · focus · HN ↗
      V4.0 Flash(-vision) in released form is "only" ~180GB of FP4 weights, so a braindead quant could work in 128GB of RAM.

      4.1 is much larger, even leaving out the PLE.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.