Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Thread
Unofficial Hacker News client; not affiliated with Y Combinator.
Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Unofficial Hacker News client; not affiliated with Y Combinator.
herf · · focus · HN ↗
Unfortunately even with my 5070ti, llama.cpp seems to be about 20-30% faster at decode, running as:
set CUDA_VISIBLE_DEVICES=0 build\bin\Release\llama-server -hf google/gemma-4-12B-it-qat-q4_0-gguf -ngl 99 --no-mmproj-offload -mg 0 -c 262144 -fa on --host 0.0.0.0
anerli · · focus · HN ↗
Currently we don't support multi-GPU setups, that is on our near-term roadmap. It saying the model is too big for that GPU might be a bug - would you be willing to open a github issue with more detail on your setup? <a href="https://github.com/magnitudedev/magnitude/issues" rel="nofollow">https://github.com/magnitudedev/magnitude/issues
As for performance, there may be some variability still depending on the model and backend. We have room for improvement for various setups that we are closing as we work out some details with our kernels and tuning system, so appreciate the data point and will look into that combination.
herf · · focus · HN ↗
would love to try it again when you have updates