‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. snehesht · · focus · HN ↗
    I tried it and it worked surprisingly well. On my machine (Nvidia 4090, 128GB DDR5, Ryzen 7950x3d) I'm getting 124 tokens per sec, thought to share it here.

    <a href="https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next" rel="nofollow">https:&#x2F;&#x2F;huggingface.co&#x2F;Qwen&#x2F;Qwen3.8-Flash-Next

    1. thatsabadlook · · focus · HN ↗
      Why is this surprisingly well? It&#x27;s 2.5x faster than anthropic models, you have data sovereignty, privacy,and that&#x27;s a strong model. Sounds like a best case scenario to me
      1. hdjrudni · · focus · HN ↗
        Not sure you understand the term &#x27;surprisingly well&#x27;. It means &#x27;better than expected&#x27;. I suspect they parent poster didn&#x27;t actually expect to get &gt;= 100 T&#x2F;s.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.