‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. ai_ja_nai · · focus · HN ↗
    I am not getting it: I see a fp2 quantized model going on a 5090 with 64GB of RAM at 90 tops with -10% accuracy over original model. How is this supportive of the claims?
    1. ai_ja_nai · · focus · HN ↗
      (64GB not VRAM, I meant) I also see people claiming fast performance on a 128GB machine, which is not exactly consumer hardware)
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.