‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

806 points · 356 comments · snehesht

  1. Tepix · · focus · HN ↗
    Q2 quantization. Not interested.
    1. CamperBob2 · · focus · HN ↗
      Larger models can tolerate Q2 quantization surprisingly well, especially if they were trained with quantization in mind. I don't know about 3.8 125B, but for example, there are 2-bit quants of Kimi K3 that exhibit strong reasoning and maintain decent coherence at longer contexts.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.