‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. kmike84 · · focus · HN ↗
    How accurate are speed estimates in the UI? I'm asking because for Qwen 3.8 (Q8) the speed numbers cited in the UI look quite poor:

      Estimated speed on your machine
      Context tokens Tokens / sec
      25 000 17
      50 000 16
      75 000 16
      262 144 12
    
    
    262K number is ok, but for lower context sizes (<128K) it's about 2x slower than the numbers I'm getting from real mtplx sessions for qwen3.8 q8 (mac m5 max).

    Is it a lack of optimizations, or incorrect numbers, or a benchmark artifact (e.g. something which is harder for spec decoding than usual agentic sessions)?

    1. francisjp · · focus · HN ↗
      To OP: great work on the release! I am generally interested in this kind of optimization work.

      Related to the post above: Similar results here M5 Max running Qwen3.8 UD-Q6-K-XL with zlab’s Dflash2 as the drafter.

      Both prefill and decode are roughly 2x faster when served from llama.cpp (b10853 or newer) than magnitude 0.2.1.

      1. anerli · · focus · HN ↗
        Thanks for pointing this out. I think we have a gap here where we may not be fully leveraging the new matmul operations available on M5+ chips, so will work that into our kernels soon and benchmark on M5 hardware.
        1. francisjp · · focus · HN ↗
          Sure thing, happy to share. That potential root cause makes sense. I bet magnitude will close the prefill gap then.

          llama.cpp had that same matmul gap (pre-fill operations) for the M5/A19 and newer silicon until that sha mentioned above.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.