‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. bythreads · · focus · HN ↗
    Ok so i took the time to benchmark this on the following on my m5 max 128gb:

    Qwen3-4B-Instruct-2507-4bit Qwen3.5-35B-A3B-4bit Qwen3.5-9B-MLX-4bit Qwen3-Reranker-0.6B-4bit Qwen3-Coder-30B-A3B-Instruct-4bit qwen2.5:0.5b

    and the results are what i kinda expected to begin with, this adds next to nothing? - also the repo was pivoted from a playwright sub assembly to this not long ago - so my conclusion - THIS MIGHT be worth some watching if you have a model where no-one!, has optimized it at all - and where it does not use anything native to your platform.

    results (averages)

    VIA rapid-mlx :8902 (MLX) Decode: 175 tok/s TTFT: 64 ms prefill (~760 tok cold): 594 ms

    Magnitude 0.2.1 (GGUF/llama.cpp+Rust) Decode: 161 tok/s TTFT: 111 ms prefill (~760 tok cold): 669 ms

    1. pbronez · · focus · HN ↗
      Rapid-MLX was my first thought too. It’s optimized for self-hosted agents on Apple silicon. It’s my current choice for self-hosted models.

      <a href="https:&#x2F;&#x2F;github.com&#x2F;raullenchai&#x2F;Rapid-MLX" rel="nofollow">https:&#x2F;&#x2F;github.com&#x2F;raullenchai&#x2F;Rapid-MLX

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.