‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. mrinterweb · · focus · HN ↗
    I'm giving this a go, and so far, this is great. I have 1x RTX 4090 (24GB VRAM) and 128GB DDR4 RAM. I am seeing > 110 tokens/sec (3 token MTP). Using 60K/260k context currently. So far the results seem at least on par with my Qwen 3.8 q5 27B (~60 TPS). I realize ultimately tokens/sec don't really mean much if the quality sucks, but I am optimistic but still paying close attention to the results.

    Working with fast local models can be great. Fast prefill, and >100 TPS is quite quick.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.