‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. nateb2022 · · focus · HN ↗
    Any source on the benchmarks/methodology besides the image? There's a ton of variance possible in llama.cpp's performance depending on how it was configured. I'd also like to see benchmarks against MLX.
    1. anerli · · focus · HN ↗
      The benchmark we cited here is a simple prose-repetition task. We put the content of Moby Dick up to 64k context in the request, and then ask it to repeat the last section.

      For llama.cpp, we try to make the comparison as fair as possible by using similar settings. No speculative decoding, default prefill batch sizes, flash attention on.

      We tried also quantizing the KV cache to 8-bit keys and 4-bit values like we do in Magnitude, but this bombed decode speed for llama.cpp in our testing. Since it seems llama.cpp did not optimize that path, we used 16-bit KV instead.

      The source for the benchmark is available here also: <a href="https:&#x2F;&#x2F;github.com&#x2F;magnitudedev&#x2F;magnitude&#x2F;tree&#x2F;main&#x2F;inference&#x2F;benchmarks&#x2F;src&#x2F;magnitude_benchmarks&#x2F;session_bench" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;magnitudedev&#x2F;magnitude&#x2F;tree&#x2F;main&#x2F;inferenc...

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.