Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Unofficial Hacker News client; not affiliated with Y Combinator.
kmike84 · · focus · HN ↗
Is it a lack of optimizations, or incorrect numbers, or a benchmark artifact (e.g. something which is harder for spec decoding than usual agentic sessions)?
francisjp · · focus · HN ↗
Related to the post above: Similar results here M5 Max running Qwen3.8 UD-Q6-K-XL with zlab’s Dflash2 as the drafter.
Both prefill and decode are roughly 2x faster when served from llama.cpp (b10853 or newer) than magnitude 0.2.1.
anerli · · focus · HN ↗
francisjp · · focus · HN ↗
llama.cpp had that same matmul gap (pre-fill operations) for the M5/A19 and newer silicon until that sha mentioned above.
anerli · · focus · HN ↗
For m5 - there may be some issue with the Metal 4 matmul hardware utilization that could be causing this to be behind here. Will look into this.