‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

805 points · 355 comments · snehesht

  1. c4pt0r · · focus · HN ↗
    100 tok/s on a 4090 could unlock entirely new use cases. What becomes possible when local inference is this fast?
    1. semireg · · focus · HN ↗
      Just like IRL: Slower and “smart enough” is better than fast and mistaken.
      1. latentsea · · focus · HN ↗
        At the end of the day all that matters is if it works for you to complete your tasks. I'm getting better results faster with this now than I was with my previous setup. So... meh?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.