Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Thread
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Loading the complete thread in the background. This saved snapshot is available now. Refresh
Unofficial Hacker News client; not affiliated with Y Combinator.
p-e-w · · focus · HN ↗
anerli · · focus · HN ↗
amirhesham · · focus · HN ↗
anerli · · focus · HN ↗
nateb2022 · · focus · HN ↗
anerli · · focus · HN ↗
For llama.cpp, we try to make the comparison as fair as possible by using similar settings. No speculative decoding, default prefill batch sizes, flash attention on.
We tried also quantizing the KV cache to 8-bit keys and 4-bit values like we do in Magnitude, but this bombed decode speed for llama.cpp in our testing. Since it seems llama.cpp did not optimize that path, we used 16-bit KV instead.
The source for the benchmark is available here also: <a href="https://github.com/magnitudedev/magnitude/tree/main/inference/benchmarks/src/magnitude_benchmarks/session_bench" rel="nofollow">https://github.com/magnitudedev/magnitude/tree/main/inferenc...
anerli · · focus · HN ↗
sgtwompwomp · · focus · HN ↗
anerli · · focus · HN ↗
It's not a coding agent running on your device optimizing the kernels, we have a system for writing kernels that can be tuned on the target device automatically. So we write the efficient high level kernel structure with tunable parameters, then it fits to whatever hardware it's actually running on.
kenzic · · focus · HN ↗
anerli · · focus · HN ↗
kenzic · · focus · HN ↗
yolandac · · focus · HN ↗
anerli · · focus · HN ↗
However we also have expert streaming on the roadmap. This will let you run mixture-of-experts models with unused experts offloaded to RAM or disk, and load them only when needed. This means you'll be able to run models that wouldn't otherwise fit in your GPU memory.
kmike84 · · focus · HN ↗
I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of he more optimized engines.
3 main failure modes I observed in the engines:
* Not using best available spec decoding * Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models) * Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline
anerli · · focus · HN ↗
Spec decoding: Models in our catalog come assigned with an assigned drafter model for speculative decoding based on the best known method and model available for that target model (support DFlash, DSpark, and DFlash2).
Using too much memory for KV cache: We use a TurboQuant-inspired quantization of KV cache to 8-bit keys and 4-bit values. This drops KV memory usage by over half and also speeds up decode. Based on long context quality benchmarking we've done it does not seem to negatively impact retrieval or coherence over long context.
Large context sizes: our KV quantization helps a lot for this, and we focus our optimizations on specifically longer-context requests since that's what most agent inference actually looks like.
skohan · · focus · HN ↗
anerli · · focus · HN ↗
All our benchmarks are open source so you can check it out here if you'd like: <a href="https://github.com/magnitudedev/magnitude/blob/main/inference/benchmarks/src/magnitude_benchmarks/fixtures/ruler.py" rel="nofollow">https://github.com/magnitudedev/magnitude/blob/main/inferenc...
kmike84 · · focus · HN ↗
Quantization is another thing. There are so many engines launched with claims about speed, but in many cases it's optimizing specific lower-quality quants. When you have enough resources, you usually want something like W8A16 + full precision KV cache working as fast as possible, not yet another W8A8 or W4A16.
In general, it seems new models are released so fast now - engines don't always have time to really polish the implementation before the next model is released
sheephess44 · · focus · HN ↗
[dead]
[deleted] · · focus · HN ↗
[deleted]
lxe · · focus · HN ↗
Then it rebuilds latest llama.cpp, grabs the PRs it finds relevant to test against, and then it performs a benchmark and finalizes the upgrade and verifies what model, variant, or even a separate finetune that we should be running.
Occasionally, it performs its own optimizations and commits, which then gets superseded by pull requests and merged code that essentially validates the model's own optimization directionality.
anerli · · focus · HN ↗
Definitely still helps to reference relevant academic work as well, or even just encouraging the agent to make bigger structural leaps, otherwise it will often get stuck working on low impact micro-optimizations.
bredren · · focus · HN ↗
genxy · · focus · HN ↗
msdz · · focus · HN ↗
Q: From my (very, very limited!) understanding, I was under the impression that part of the “inference engine inertia” is model- or at least architecture-specific code for most, if not each new open-weight model coming out. Assuming I got that right, do you support, or plan on supporting, everything vLLM/llama.cpp do, such that Magnitude becomes a drop-in replacement for as many (economically/pareto-viable) models as possible, or do you want to focus on the best possible support for only a select few models/classes of models?
anerli · · focus · HN ↗
We plan to support any model architecture that we believe is somewhere along or close to the pareto frontier. There's some model families that are outdated or more niche that we don't necessarily want to put our focus into.
msdz · · focus · HN ↗
herf · · focus · HN ↗
Unfortunately even with my 5070ti, llama.cpp seems to be about 20-30% faster at decode, running as:
set CUDA_VISIBLE_DEVICES=0 build\bin\Release\llama-server -hf google/gemma-4-12B-it-qat-q4_0-gguf -ngl 99 --no-mmproj-offload -mg 0 -c 262144 -fa on --host 0.0.0.0
anerli · · focus · HN ↗
Currently we don't support multi-GPU setups, that is on our near-term roadmap. It saying the model is too big for that GPU might be a bug - would you be willing to open a github issue with more detail on your setup? <a href="https://github.com/magnitudedev/magnitude/issues" rel="nofollow">https://github.com/magnitudedev/magnitude/issues
As for performance, there may be some variability still depending on the model and backend. We have room for improvement for various setups that we are closing as we work out some details with our kernels and tuning system, so appreciate the data point and will look into that combination.
herf · · focus · HN ↗
would love to try it again when you have updates
sheephess44 · · focus · HN ↗
[dead]
teabee89 · · focus · HN ↗
anerli · · focus · HN ↗
Magnitude is optimized for maximum single-session performance and memory efficiency - so we should be more performant for local inference use cases.
hypercube33 · · focus · HN ↗
anerli · · focus · HN ↗
Regarding model variants - our catalog includes different quantizations, and automatically assesses these against your hardware to determine which ones will fit in your memory and how fast they will run. This lets you pick a model to download based on your desired speed/intelligence tradeoff.
skohan · · focus · HN ↗
anerli · · focus · HN ↗
hypercube33 · · focus · HN ↗
hypercube33 · · focus · HN ↗
Install is easy. I like the ability to find models for my system with estimated performance.
Cons: I cant re-use models for other systems like LM Studio I have saved to a folder on my system; it wants to re-download 40GB models.
mncharity · · focus · HN ↗
External/policy-based throttling for temperature control. Unthrottled, my laptop bottom goes skin-burn hot. But fixed compute caps can have non-linearly dreadful performance impacts in particular cases. Plan is a runtime knob, to replace manual limits-kludgery.
I'll use models which barely fit in VRAM+RAM, and are order-1 tok/s slow. So tool call step overhead can be painful - a world where `ls` costs tens of seconds. Plan is blending harness plugins with inference loop, for "no, don't stop - I already have the call result for you - just keep going" (and also some logit games).
anerli · · focus · HN ↗
<a href="https://github.com/magnitudedev/magnitude/issues" rel="nofollow">https://github.com/magnitudedev/magnitude/issues
aidiscoverywire · · focus · HN ↗
[dead]
paulgerhardt · · focus · HN ↗
anerli · · focus · HN ↗
<a href="https://github.com/magnitudedev/magnitude/issues" rel="nofollow">https://github.com/magnitudedev/magnitude/issues
Let me know if you keep running into problems for some reason
cedricd · · focus · HN ↗
anerli · · focus · HN ↗
Could you share your hardware and OS details to help us identify what might be the issue here?
There's also a github issue open on this topic if you want to leave a comment there: <a href="https://github.com/magnitudedev/magnitude/issues/142" rel="nofollow">https://github.com/magnitudedev/magnitude/issues/142
c7b · · focus · HN ↗
And a more general question: does your engine detect and optimize / suggest a good load split for custom setups like multiple (possibly different) GPUs, eGPUs,...? Because if all you have is a stock major system like a Mac or DGX Spark, that's all you're going to care about, and there are a lot of highly optimized single-hardware engines out there that will be hard to beat in the long run. Something that automatically adapts to custom systems that don't have their own subreddits could really fill a gap.
anerli · · focus · HN ↗
Qwen3.8-Flash-Next support will also be added very soon.
Taking full advantage of all the hardware on your machine in the most performant way possible is the overall goal of the inference engine. This includes a lot of what you're describing. We want to map out the full hardware topology of your system (one or more GPUs, CPU, memory), and compile a combination of kernels to serve a given model optimally across that stack, allocating different parts of the workload wherever it fits best.
Currently we're writing tunable kernels that optimize themselves for one device, but we're working on a kernel compiler that will be able to compile and distribute kernels across any number of devices in a system.
MaxikCZ · · focus · HN ↗
Can it do all the shenanigans that allows to run qwen flash on 12GB vram over 40 toks like people seems to be getting in this thread?: <a href="https://www.reddit.com/r/LocalLLaMA/comments/1wp7zyb/qwen38flashnext_on_12gb_vram_65_tokens_per_second/" rel="nofollow">https://www.reddit.com/r/LocalLLaMA/comments/1wp7zyb/qwen38f...
anerli · · focus · HN ↗
Our plan to enable running bigger models on less GPU memory in a way that'll remain productive is expert streaming. This will let you offload experts for MoE models to RAM or disk, and load them when needed. This can have some performance tradeoff, but is lossless.
MaxikCZ · · focus · HN ↗
ls-a · · focus · HN ↗
[dead]
kmike84 · · focus · HN ↗
Estimated speed on your machine Context tokens Tokens / sec 25 000 17 50 000 16 75 000 16 262 144 12
262K number is ok, but for lower context sizes (<128K) it's about 2x slower than the numbers I'm getting from real mtplx sessions for qwen3.8 q8 (mac m5 max).
francisjp · · focus · HN ↗
Both prefill and decode are roughly 2x faster when served from llama.cpp than magnitude 0.2.1.
anerli · · focus · HN ↗
francisjp · · focus · HN ↗
llama.cpp had that same matmul gap (pre-fill operations) for the M5/A19 and newer silicon until that sha mentioned above.
anerli · · focus · HN ↗
For m5 - there may be some issue with the Metal 4 matmul hardware utilization that could be causing this to be behind here. Will look into this.
singh_abinashi · · focus · HN ↗
andrethegiant · · focus · HN ↗
happybox2016 · · focus · HN ↗
anerli · · focus · HN ↗
llama.cpp does not saturate memory bandwidth for single-stream tok/s, and for long context and batching, our quantized KV and associated decode kernels allow us to reduce the effective bandwidth needed, and surpass llama.cpp significantly in decode speeds.
williamse · · focus · HN ↗
Zetaphor · · focus · HN ↗
anerli · · focus · HN ↗
digitaltrees · · focus · HN ↗
I am building propelcompute.com an open router for private hardware and experimented with exo labs to run large models on for Mac studios and plan on doing the same with nvidia and amd. Id love to integrate your inference engine into the system but built gpu is critical.
anerli · · focus · HN ↗
digitaltrees · · focus · HN ↗
taylorhou · · focus · HN ↗
The autotuner is the real thing - kernel-level search, config budget split by measured time share, winners cached per device/toolchain. Rotating resident weights across layers to dodge the hot-cache trap is a nice touch.
One question on model ranking: the fit scores look like predictions built from cost constants measured on a single M4 Max, not per-device measurements. How do you rank models across genuinely mixed hardware? Does the estimator improve from actual runs over time? That gap between predicted and measured fit is what eats mixed machine fleets alive.
Shameless plug since magnitude's goal is highly relevant to what i'm working on: teale.com - distributed inference across fleets of macs. If you're running local models on more than one box, check it out with your agent!
anerli · · focus · HN ↗
ContinuityLab · · focus · HN ↗
[dead]
chzblck · · focus · HN ↗
On a 64gb Ram and 5080 machine the biggest model suggested was Qwen 9B
can hit 90+ tps on the MoE 35b but mag thinks it wont fit.
anerli · · focus · HN ↗
The MoE 35b might be tight though unless you were to go below 4-bit. Could you share the quant you used when you ran this on that 5080 before? Our catalog only contains models down to 4-bit because we find that thinking, tool calling, and overall capabilities start to suffer at lower fidelity.
Feel free also to create a GitHub issue with more details and we can take a closer look.
lin7c · · focus · HN ↗
thoughtpeddler · · focus · HN ↗
anerli · · focus · HN ↗
So it's more about making those GPU kernels perform the required memory move and arithmetic operations as close to the theoretical optimum as possible.
loclol101 · · focus · HN ↗
anerli · · focus · HN ↗
thoughtpeddler · · focus · HN ↗
--
[0] <a href="https://developer.apple.com/videos/play/wwdc2026/326/" rel="nofollow">https://developer.apple.com/videos/play/wwdc2026/326/
anerli · · focus · HN ↗
thoughtpeddler · · focus · HN ↗
tinykit · · focus · HN ↗
[dead]
sebastienburel · · focus · HN ↗
Second, more important for agents: decode speed is rarely what hurts. It's resending the same system prompt plus tool schemas every turn. Does self-optimizing cover prefix cache reuse across requests, or is it kernel and layout tuning only?
And is the endpoint OpenAI-compatible? My runtime already talks to llama.cpp and LM Studio through that wrapper, so drop-in is the difference between trying it tonight and not.
bythreads · · focus · HN ↗
anerli · · focus · HN ↗
Prefix cache is re-used with a prefix tree structure for maximal re-use across sessions sharing prompts.
The endpoint is standard OpenAI compatible chat completions.
aitoolcrux · · focus · HN ↗
The hard part isn't the optimization itself—it's measuring whether a change actually helps across the long tail of request patterns. Most inference benchmarks show huge gains on popular workloads, but production traffic has a fat tail where naive optimizations hurt latency.
Interested in how you handle regression detection when the optimizer changes between requests.
bythreads · · focus · HN ↗
Qwen3-4B-Instruct-2507-4bit Qwen3.5-35B-A3B-4bit Qwen3.5-9B-MLX-4bit Qwen3-Reranker-0.6B-4bit Qwen3-Coder-30B-A3B-Instruct-4bit qwen2.5:0.5b
and the results are what i kinda expected to begin with, this adds next to nothing? - also the repo was pivoted from a playwright sub assembly to this not long ago - so my conclusion - THIS MIGHT be worth some watching if you have a model where no-one!, has optimized it at all - and where it does not use anything native to your platform.
results (averages)
VIA rapid-mlx :8902 (MLX) Decode: 175 tok/s TTFT: 64 ms prefill (~760 tok cold): 594 ms
Magnitude 0.2.1 (GGUF/llama.cpp+Rust) Decode: 161 tok/s TTFT: 111 ms prefill (~760 tok cold): 669 ms
anerli · · focus · HN ↗
pbronez · · focus · HN ↗
<a href="https://github.com/raullenchai/Rapid-MLX" rel="nofollow">https://github.com/raullenchai/Rapid-MLX
larodi · · focus · HN ↗
theParadox42 · · focus · HN ↗
anerli · · focus · HN ↗
alescalaios · · focus · HN ↗
[dead]
NKosmatos · · focus · HN ↗
anerli · · focus · HN ↗
Plus in the near future, we'll add ways to automatically utilize your RAM as well (such as expert-streaming).
Our goal is to make the best use of whatever hardware you have, even if its not high end!
mrtsepelev · · focus · HN ↗
anerli · · focus · HN ↗
eventuallyworth · · focus · HN ↗
[dead]
nullbio · · focus · HN ↗
pxtail · · focus · HN ↗
pxtail · · focus · HN ↗
karlkloss · · focus · HN ↗
anerli · · focus · HN ↗
<a href="https://github.com/magnitudedev/magnitude/issues" rel="nofollow">https://github.com/magnitudedev/magnitude/issues
[deleted] · · focus · HN ↗
[deleted]
steinvakt2 · · focus · HN ↗
anerli · · focus · HN ↗
malshe · · focus · HN ↗
anerli · · focus · HN ↗
openamer · · focus · HN ↗
vivzkestrel · · focus · HN ↗