Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Unofficial Hacker News client; not affiliated with Y Combinator.
sebastienburel · · focus · HN ↗
Second, more important for agents: decode speed is rarely what hurts. It's resending the same system prompt plus tool schemas every turn. Does self-optimizing cover prefix cache reuse across requests, or is it kernel and layout tuning only?
And is the endpoint OpenAI-compatible? My runtime already talks to llama.cpp and LM Studio through that wrapper, so drop-in is the difference between trying it tonight and not.
anerli · · focus · HN ↗
Prefix cache is re-used with a prefix tree structure for maximal re-use across sessions sharing prompts.
The endpoint is standard OpenAI compatible chat completions.