‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. happybox2016 · · focus · HN ↗
    2x llama.cpp" on what, an M3 Max? llama.cpp's metal kernels already saturate memory bandwidth. Real agent bottleneck isn't single-stream tok/s — it's KV cache for 5+ concurrent 128k contexts on 24GB VRAM. Who's actually running multi-agent locally? A) Single session only B) 2-3 agents C) 5+ agents D) Gave up,
    1. anerli · · focus · HN ↗
      On our benchmarks we approach 2x decode speeds on a variety of Mac hardware (tested most on M4 Pro and Max).

      llama.cpp does not saturate memory bandwidth for single-stream tok/s, and for long context and batching, our quantized KV and associated decode kernels allow us to reduce the effective bandwidth needed, and surpass llama.cpp significantly in decode speeds.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.