‹ BackHN Continuity

Thread

Run Qwen 3.8 Flash Next (125B) on consumer hardware (RTX 4090) at 100T/s

741 points · 330 comments · snehesht

  1. lxe · · focus · HN ↗
    Why isn't this type of expert caching in the native llama.cpp yet? Why do we need a separate codebase?
    1. parsimo2010 · · focus · HN ↗
      One reason is that the llama.cpp team (GGML) has strict requirements that a human must understand the code they are contributing. If a project is fully vibe coded they can’t contribute. So a lot of projects where an AI went and coded a bunch of custom kernels to increase speed are left to their own devices.

      I think this is a fine behavior. We can have upstream purists that are strict gatekeepers but don’t get in the way of downstream forks. Debian has some this in the Linux landscape for a long time, and it has enabled Ubuntu, Mint, etc. to flourish without compromising themselves.

      1. Loquebantur · · focus · HN ↗
        "A human pretends to understand it" signifies what exactly?

        What you really mean is, the core team there doesn't want to lose control.

        Which isn't really predicated on contributions not being "vibe coded" or whatever.

        When quality is the problem, you need to be able to make your standards explicit, or you're just gatekeeping irrationally.

        1. anamexis · · focus · HN ↗
          They do make their standards explicit: <a href="https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;CONTRIBUTING.md#ai-usage-policy" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;ggml-org&#x2F;llama.cpp&#x2F;blob&#x2F;master&#x2F;CONTRIBUTI...

          What part do you think is irrational gatekeeping?

          1. rfgplk · · focus · HN ↗
            &gt; A proper code review usually takes something like one hour per 200-400 LOC and you should be spending at least that much time on code review alone.

            Not only is this not enforceable (how do you enforce how long someone spent working on a codebase on their own local machine?) the metric is severely off which instantly makes me question the competence of the llama.cpp dev team. You can easily review 10-100x that in an hour, even if you&#x27;re being super pedantic about it.

            I also just ran _one_ of their files (with include deps) through Astra and it detected &gt;100 vulnerabilities&#x2F;correctness errors (with over 10 outright UB&#x2F;memory corruption issues). It&#x27;s actually outright shocking.

            1. kube-system · · focus · HN ↗
              What you have quoted is a sentence elaborating on the requirements listed in the document. This is provided to help you better understand why the requirements exist and the goal they are trying to accomplish.

              &gt; should

              <a href="https:&#x2F;&#x2F;www.rfc-editor.org&#x2F;info&#x2F;rfc2119&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.rfc-editor.org&#x2F;info&#x2F;rfc2119&#x2F;

              The reason you SHOULD take that time to read the output is because you must read it to understand it.

              And the way this is enforced is explicitly called out in the document (and again in more detail in the linked AGENTS.md): the maintainers may ask you to explain it.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.