‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. happybox2016 · · focus · HN ↗
    2x llama.cpp" on what, an M3 Max? llama.cpp's metal kernels already saturate memory bandwidth. Real agent bottleneck isn't single-stream tok/s — it's KV cache for 5+ concurrent 128k contexts on 24GB VRAM. Who's actually running multi-agent locally? A) Single session only B) 2-3 agents C) 5+ agents D) Gave up,
    1. williamse · · focus · HN ↗
      B, 2-3 agents. The KV cache framing is the right one. Single-stream tok/s is what shows up in benchmarks but it's not what actually hurts when agents are sleeping between tool calls and waking up needing their full context. The question I'd want answered about an engine like this is how it handles partially-cold contexts, because agent sessions aren't uniform sustained reads, they're bursty and interleaved.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.