‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. kmike84 · · focus · HN ↗
    How accurate are speed estimates in the UI? I'm asking because for Qwen 3.8 (Q8) the speed numbers cited in the UI look quite poor:

      Estimated speed on your machine
      Context tokens Tokens / sec
      25 000 17
      50 000 16
      75 000 16
      262 144 12
    
    
    262K number is ok, but for lower context sizes (<128K) it's about 2x slower than the numbers I'm getting from real mtplx sessions for qwen3.8 q8 (mac m5 max).

    Is it a lack of optimizations, or incorrect numbers, or a benchmark artifact (e.g. something which is harder for spec decoding than usual agentic sessions)?

    1. francisjp · · focus · HN ↗
      To OP: great work on the release! I am generally interested in this kind of optimization work.

      Related to the post above: Similar results here M5 Max running Qwen3.8 UD-Q6-K-XL with zlab’s Dflash2 as the drafter.

      Both prefill and decode are roughly 2x faster when served from llama.cpp (b10853 or newer) than magnitude 0.2.1.

      1. anerli · · focus · HN ↗
        Thanks for pointing this out. I think we have a gap here where we may not be fully leveraging the new matmul operations available on M5+ chips, so will work that into our kernels soon and benchmark on M5 hardware.
        1. francisjp · · focus · HN ↗
          Sure thing, happy to share. That potential root cause makes sense. I bet magnitude will close the prefill gap then.

          llama.cpp had that same matmul gap (pre-fill operations) for the M5/A19 and newer silicon until that sha mentioned above.

    2. anerli · · focus · HN ↗
      These numbers are just estimates based on your hardware and may differ from actual performance. It's hard to get an accurate measurement until it's actually downloaded and running. They also don't account for gains from speculative decoding. Working on changes to make this more clear.

      For m5 - there may be some issue with the Metal 4 matmul hardware utilization that could be causing this to be behind here. Will look into this.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.