‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. lin7c · · focus · HN ↗
    One thing I'd want to see in the evals is per-turn latency across a full agent trajectory, not just end-to-end time. In my experience the workload flips mid-run: early turns are prefill-heavy (big system prompt, tool schemas), late turns are short decodes against a huge KV cache, so a config that's optimal for turn one can be badly wrong by turn thirty. The self-tuning idea is interesting, but I'm curious whether the tuning happens per-request or per-trajectory. With prefix-cached tool schemas the win should compound; without it you're re-solving the same optimization problem every call.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.