‹ BackHN Continuity

Thread

Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents

194 points · 99 comments · anerli

  1. MaxikCZ · · focus · HN ↗
    if fully custom compiler would find best settings for given setup, upload the setup to mothership and allow new peers to download it as good starting point.

    Can it do all the shenanigans that allows to run qwen flash on 12GB vram over 40 toks like people seems to be getting in this thread?: <a href="https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1wp7zyb&#x2F;qwen38flashnext_on_12gb_vram_65_tokens_per_second&#x2F;" rel="nofollow">https:&#x2F;&#x2F;www.reddit.com&#x2F;r&#x2F;LocalLLaMA&#x2F;comments&#x2F;1wp7zyb&#x2F;qwen38f...

    1. anerli · · focus · HN ↗
      While it&#x27;s great to see tok&#x2F;s go up as high as possible, I think it&#x27;s important to consider the actual usability of these models when you quantize down to something like 2-bit. From what we&#x27;ve tested it seems like going below 4-bit quickly leads to serious issues with thinking, tool calls, and overall model coherence.

      Our plan to enable running bigger models on less GPU memory in a way that&#x27;ll remain productive is expert streaming. This will let you offload experts for MoE models to RAM or disk, and load them when needed. This can have some performance tradeoff, but is lossless.

      1. MaxikCZ · · focus · HN ↗
        I totally get that, Its just, regardless of how bit-quantized it is, they are still pulling 40t&#x2F;s from model sitting mostly in RAM instead of VRAM. If I understand correctly they split the network parts very deliberately between VRAM and RAM, and I wonder if your program, of which main feature is &quot;get most of your hardware&quot; is capable of similar feats, or if that performace is still locked for those willing to spend days experimenting manually.
Open on Hacker News to reply ↗

Unofficial Hacker News client; not affiliated with Y Combinator.