‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. kmike84 · · focus · HN ↗
    This seems to be a good idea. However, beating llama.cpp on speed is a low bar :)

    I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of the more optimized engines.

    3 main failure modes I observed in the engines:

    * Not using best available spec decoding

    * Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models)

    * Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline

    1. anerli · · focus · HN ↗
      Yeah these are all things that we directly tackle!

      Spec decoding: Models in our catalog come assigned with an assigned drafter model for speculative decoding based on the best known method and model available for that target model (support DFlash, DSpark, and DFlash2).

      Using too much memory for KV cache: We use a TurboQuant-inspired quantization of KV cache to 8-bit keys and 4-bit values. This drops KV memory usage by over half and also speeds up decode. Based on long context quality benchmarking we've done it does not seem to negatively impact retrieval or coherence over long context.

      Large context sizes: our KV quantization helps a lot for this, and we focus our optimizations on specifically longer-context requests since that's what most agent inference actually looks like.

      1. kmike84 · · focus · HN ↗
        I think the specific issue I had was due to lack of proper support for ds4 compressed KV cache, not about KV cache quantization. It was like 50GB instead of 5GB for context, and it wasn't fixed for weeks (I haven't checked if it's fixed now - hopefully it is).

        Quantization is another thing. There are so many engines launched with claims about speed, but in many cases it's optimizing specific lower-quality quants. When you have enough resources, you usually want something like W8A16 + full precision KV cache working as fast as possible, not yet another W8A8 or W4A16.

        In general, it seems new models are released so fast now - engines don't always have time to really polish the implementation before the next model is released

        1. sheephess44 · · focus · HN ↗

          [dead]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.