‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. chzblck · · focus · HN ↗
    Sounds interesting would love to test it out but here's what I got when I first tried to get some models downloaded.

    On a 64gb Ram and 5080 machine the biggest model suggested was Qwen 9B

    can hit 90+ tps on the MoE 35b but mag thinks it wont fit.

    1. anerli · · focus · HN ↗
      Yeah that doesn't sound quite right. Given that the 5080 has 16 GB of VRAM I would have expected a few more options, for example Gemma 12B, to at least show up as available. Do any bigger models show as available or was that specifically the recommendation?

      The MoE 35b might be tight though unless you were to go below 4-bit. Could you share the quant you used when you ran this on that 5080 before? Our catalog only contains models down to 4-bit because we find that thinking, tool calling, and overall capabilities start to suffer at lower fidelity.

      Feel free also to create a GitHub issue with more details and we can take a closer look.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.