‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. lxe · · focus · HN ↗
    Why isn't this type of expert caching in the native llama.cpp yet? Why do we need a separate codebase?
    1. parsimo2010 · · focus · HN ↗
      One reason is that the llama.cpp team (GGML) has strict requirements that a human must understand the code they are contributing. If a project is fully vibe coded they can’t contribute. So a lot of projects where an AI went and coded a bunch of custom kernels to increase speed are left to their own devices.

      I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.

      1. yieldcrv · · focus · HN ↗
        ditch llama.cpp, its for 2024 and stuck in 2024
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.