‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

806 points · 356 comments · snehesht

  1. gdevenyi · · focus · HN ↗
    I had this working with the FreeToken inference engine a month ago when they launched.

    <a href="https:&#x2F;&#x2F;github.com&#x2F;FlashML-org&#x2F;FreeToken" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;FlashML-org&#x2F;FreeToken

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.