‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. digitaltrees · · focus · HN ↗
    Do you support splitting models across devices so larger models can run on clusters?

    I am building propelcompute.com an open router for private hardware and experimented with exo labs to run large models on for Mac studios and plan on doing the same with nvidia and amd. Id love to integrate your inference engine into the system but built gpu is critical.

    1. anerli · · focus · HN ↗
      The goal of the inference engine is to make the best use of whatever hardware you have to run models performantly and let people run bigger models. At first this will include using all the hardware on a given machine optimally. Eventually we also want to support interconnect between multiple machines to enable running bigger models!
      1. digitaltrees · · focus · HN ↗
        Let me know if youre interested in a collaboration then. I am working on a custom mlx sharding system.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.