ldraw-nova: AI agents build LEGO models by writing Python geometry
In ldraw-nova, models write a plan and a Python generator instead of raw coordinates, but nothing checks physics or stability.
Magnitude's 2x speedup is decode, but at 64K context an agent turn is mostly prefill, so the whole turn gets only about 12% faster.

"Up to 2x faster than llama.cpp" was the headline on Magnitude's Launch HN this week. The 2x is decode. Prefill moved 9%.
Magnitude is an open source local inference engine for agents. It tunes its GPU kernels on your own machine, about a minute per model.
Their Mac numbers, Qwen 35B at 64K context:
I ran the math for one cold agent turn. Reading 64K tokens takes over two minutes on both engines. Writing a 500 token answer drops from 17 seconds to 9. The turn takes about 12% less time, not half.
An agent turn is mostly reading: tool output, diffs, whole files. That's prefill.
Before you switch engines, measure:
Decode sells the demo. Prefill decides the wait.