‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. mrtsepelev · · focus · HN ↗
    Congrats on launch! Tried it on the gemma-4-26b-qat-4bit model. Was indeed faster on token generation then on oMLX (82.8 tok/s vs 76.5 tok/s), but the prefill time was ~2.6x slower (709 tok/s vs 1843 tok/s). Don’t use any acceleration on the oMLX. Macbook M5 Pro, 48 gb
    1. anerli · · focus · HN ↗
      Hey, yeah this is a known issue on M5+ macs. We are working on a patch so that our kernels use that hardware acceleration path. This should make prefill faster than MLX-based engines and boost decode a bit more for that hardware!
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.