Magnitude's self-optimizing engine runs open models 2x faster than llama.cpp
Launch HN: Magnitude (YC S25) – Self-optimizing inference engine for agents
Magnitude is an open-source inference engine that compiles and tunes its kernels on your device before a model runs, delivering up to 2x faster performance than llama.cpp—92% faster decode on Apple Silicon and 19% on NVIDIA CUDA. It connects to agents like Pi, OpenCode, and Codex with one click, uses 27% less memory per agent, and works on any Apple Silicon, NVIDIA, or AMD GPU, or just a CPU. Free, private, and Apache 2.0.
They ship kernels precompiled for broad classes of hardware. Magnitude compiles and tunes its kernels on your actual device before a model runs, so they fit your exact chip.
- kmike84
How accurate are speed estimates in the UI? I'm asking because for Qwen 3.8 (Q8) the speed numbers cited in the UI look quite poor:
Estimated speed on your machine
Context tokens Tokens / sec
25 000 17
50 000 16
75 000 16
262 144 12
262K number is ok, but for lower context sizes (<128K) it's about 2x slower than the numbers I'm getting from real mtplx sessions for qwen3.8 q8 (mac m5 max).
Is it a lack of optimizations, or incorrect numbers, or a benchmark artifact (e.g. something which is harder for spec decoding than usual agentic sessions)?
- happybox2016
2x llama.cpp" on what, an M3 Max? llama.cpp's metal kernels already saturate memory bandwidth. Real agent bottleneck isn't single-stream tok/s — it's KV cache for 5+ concurrent 128k contexts on 24GB VRAM. Who's actually running multi-agent locally? A) Single session only B) 2-3 agents C) 5+ agents D) Gave up,
- lxe
On my local inference box I have a perpetual codex thread open in my llama.cpp checkout that I periodically ask to take a look at currently pending llama.cpp PRs, do some research on latest MTP, Dflash and other prediction or attention optimizations, do research on the latest model quants and finetunes, take a look at localLlama Reddit threads and just do essentially a sweep of the frontier.
Then it rebuilds latest llama.cpp, grabs the PRs it finds relevant to test against, and then it performs a benchmark and finalizes the upgrade and verifies what model, variant, or even a separate finetune that we should be running.
Occasionally, it performs its own optimizations and commits, which then gets superseded by pull requests and merged code that essentially validates the model's own optimization directionality.
- kmike84
This seems to be a good idea. However, beating llama.cpp on speed is a low bar :)
I found it to be a good baseline, but at least on Mac there was always something way faster, and/or with better memory requirements - like you said, ds4, omlx, mtplx, etc. It seems if you use local LLMs for real, there is very little reason not to use one of the more optimized engines.
3 main failure modes I observed in the engines:
* Not using best available spec decoding
* Using too much VRAM for KV cache (e.g. KV cache used to take almost nothing in ds4, but huge amount of VRAM on unsloth/llama.cpp for deepseek models)
* Degraded performance at large context sizes - benchmarks at 4K or 32K are awesome, but at realistic 100-200K it's slower than some stupid baseline
- singh_abinashi
Curious how the evals for this work on real agent workloads versus synthetic benchmarks. In my experience, agent cost and latency profiles change a lot once there's a tool-use loop involved, because the token distribution gets much burstier than a single prompt. Did you evaluate on multi-step tool-calling traces, or mostly single-turn?