‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. nialv7 · · focus · HN ↗
    There are so many AI generated inference engine for local models now, each of them are generally narrower but they are all faster than llama.cpp. Maybe llama.cpp needs to rethink their strategies...
    1. mkatx · · focus · HN ↗
      Does anything else support Pascal gpu's though?
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.