‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. kmike84 · · focus · HN ↗
    This seems to be a good idea. However, beating llama.cpp on speed is a low bar :)

    I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of the more optimized engines.

    3 main failure modes I observed in the engines:

    * Not using best available spec decoding

    * Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models)

    * Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline

    1. [deleted] · · focus · HN ↗

      [deleted]

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.