ldraw-nova: AI agents build LEGO models by writing Python geometry
In ldraw-nova, models write a plan and a Python generator instead of raw coordinates, but nothing checks physics or stability.
Reasoning made Jeeves five points more accurate but pushed p90 latency to 17 s; reasoning only below 0.9 confidence keeps most of the gain.

A 9B classifier got five points more accurate once it could think before answering. Its median response went from 0.3 seconds to 3.3.
PostHog open-sourced Jeeves today, a small decision model for yes/no, multiple-choice and rating questions. The kind of call behind ticket routing or content flags.
The tail is where it hurts. With full reasoning, p90 latency on one H100 reaches 17 seconds.
The middle setting is worth copying. The model answers without thinking first and only reasons, with a capped chain, when its confidence is below 0.9. On their dev set that keeps most of the gain, 0.806 against 0.825, at a 2 second median and under 6 at p90.
Before you add reasoning to a classifier:
Reasoning is a budget. Spend it on the uncertain cases.