‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. yolandac · · focus · HN ↗
    does it allow us to run larger models that weren't possible before?
    1. anerli · · focus · HN ↗
      Right now, since we use less memory for KV, you have more room for model weights when you're running longer sessions.

      However we also have expert streaming on the roadmap. This will let you run mixture-of-experts models with unused experts offloaded to RAM or disk, and load them only when needed. This means you'll be able to run models that wouldn't otherwise fit in your GPU memory.

Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.