‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

793 points · 354 comments · snehesht

  1. lxe · · focus · HN ↗
    Why isn't this type of expert caching in the native llama.cpp yet? Why do we need a separate codebase?
    1. hgoel · · focus · HN ↗
      The "mainstream" inference engines are notoriously slow to integrate this stuff, to an extent understandably given the complexity of ensuring numerical accuracy alongside supporting a wide array of systems and models. Part of it is that not everyone is willing to bring what they develop into a pull request because they vibe coded it and don't care to deal with whatever quality requirements the more well known inference engines have.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.