Launch HN: Magnitude (YC S25) ships a self-optimizing local inference engine for agent workloads
- Magnitude, a YC S25 startup, launched an inference engine that writes kernels with tunable parameters and fits them to the target hardware, a one-time tuning pass of about one minute per downloaded model.
- A TurboQuant-inspired KV cache quantization to 8-bit keys and 4-bit values cuts KV memory by more than half and speeds decode; the team says long-context retrieval and coherence hold up.
- The headline benchmark against llama.cpp is a Moby Dick prose-repetition task at up to 64k context, where llama.cpp ran 16-bit KV because 8-bit keys and 4-bit values tanked its decode speed in Magnitude's testing.
- The business model is an inference cloud for hybrid workloads billed per token, routing between local and cloud models without breaking the prefix cache. Nothing is built yet; the repo shows 5.7k stars and 397 forks.
- Roadmap items include expert streaming to offload unused mixture-of-experts weights to RAM or disk, and catalog models that ship with assigned drafter models for speculative decoding (DFlash, DSpark, DFlash2).
Hacker News opinions
What's the business model here? Everyone is shipping local inference engines and nobody explains how they make money.
We think workloads end up hybrid. Consumer hardware handles a lot locally, cloud models for the harder tasks, and we want switching between the two to not break your prefix cache. We'll charge per token on our inference cloud once it exists.
Any source for the benchmark besides the chart? llama.cpp performance swings wildly depending on how you configure it, and I'd also like to see numbers against MLX.
It's a prose-repetition task with Moby Dick up to 64k context. We tried 8-bit keys and 4-bit values on llama.cpp the way we do in Magnitude and decode speed bombed, so we used 16-bit KV there. No speculative decoding, default prefill batch sizes, flash attention on. Source is in the repo.
On MLX we've done rough benchmarking and we're ahead of every MLX-based engine we've compared to so far. More in-depth numbers coming soon.
Is this Wafer.ai but for local models, a coding agent tuning the kernels so the model keeps getting faster? That would be compelling.
The common thread with Wafer is performant inference for agent workloads. It is not a coding agent running on your device optimizing kernels. We write the high-level kernel structure with tunable parameters, then it fits to whatever hardware it actually runs on.
How long does the tuning take on something like an M3 MacBook Pro?
About a minute, one time, whenever you download a new model. That is enough to hit the point where tuning longer asymptotes, though it varies a little by hardware.
Does this let me run larger models that weren't possible before?
Since we use less memory for KV, you get more room for weights in long sessions. Expert streaming is on the roadmap too, which offloads unused MoE experts to RAM or disk so you can run models that would not otherwise fit in GPU memory.
Beating llama.cpp on speed is a low bar. The three failure modes I see: no speculative decoding, way too much VRAM for KV cache, and performance falling off a cliff past 32k even when 4K benchmarks look great.
We tackle all three. Catalog models come with an assigned drafter for speculative decoding (DFlash, DSpark, DFlash2), KV is quantized to 8-bit keys and 4-bit values, and we optimize specifically for the long-context requests that agent inference actually is.
I keep a Codex thread open in my llama.cpp checkout that sweeps pending PRs, benchmarks them, and occasionally commits its own optimizations that later get validated by merged upstream code.
We lean on coding agents hard for our kernels. Since it is highly verifiable and slow to measure, we leave several running on different model architectures. Left alone they get stuck on low-impact micro-optimizations, so pushing them toward bigger structural leaps helps.
Do you plan to be a drop-in replacement for everything vLLM and llama.cpp support, or focus on getting a few model classes really fast?
Models in the same family share an architecture, so you optimize once and new versions reuse the same kernels, and some algorithms carry across families. We'll support anything on or near the pareto frontier, not outdated or niche families.
My two 16GB NVIDIA cards get detected as four GPUs, most models over 8GB are rejected, and it seems to run on the 5070ti alone. Even there llama.cpp decodes 20 to 30 percent faster for me.