‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. hypercube33 · · focus · HN ↗
    From your description looks like this isn't for AMD or Strix Halo at all? Also one of the things I'm not sure of but definitely plays a huge factor is the variant of the model you download - how does this help select the fastest version for your specific hardware / context size?
    1. anerli · · focus · HN ↗
      We support Vulkan as well, we just didn't mention it in the benchmark. When AMD or Strix Halo is detected the engine will use Vulkan.

      Regarding model variants - our catalog includes different quantizations, and automatically assesses these against your hardware to determine which ones will fit in your memory and how fast they will run. This lets you pick a model to download based on your desired speed/intelligence tradeoff.

      1. skohan · · focus · HN ↗
        Do you have any plans to support ROCm?
        1. anerli · · focus · HN ↗
          We are actively benchmarking our Vulkan kernels to ROCm implementations in other engines to ensure that we can reach the performance ceiling with them. Vulkan is much more portable and also works on non-AMD hardware even though it can be more awkward to write kernels for. If we find that Vulkan is not sufficient for reaching the same performance as ROCm, we'll consider adding it as a backend
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.