Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Unofficial Hacker News client; not affiliated with Y Combinator.
MaxikCZ · · focus · HN ↗
Can it do all the shenanigans that allows to run qwen flash on 12GB vram over 40 toks like people seems to be getting in this thread?: <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wp7zyb/qwen38flashnext_on_12gb_vram_65_tokens_per_second/" rel="nofollow">https://www.reddit.com/r/LocalLLaMA/comments/1wp7zyb/qwen38f...
anerli · · focus · HN ↗
Our plan to enable running bigger models on less GPU memory in a way that'll remain productive is expert streaming. This will let you offload experts for MoE models to RAM or disk, and load them when needed. This can have some performance tradeoff, but is lossless.
MaxikCZ · · focus · HN ↗