ldraw-nova: AI agents build LEGO models by writing Python geometry
In ldraw-nova, models write a plan and a Python generator instead of raw coordinates, but nothing checks physics or stability.
livenerf's validation found accuracy noise bigger than the real change, while output token counts clearly showed effort shifts.

Every few weeks someone swears a model got worse after launch. Almost nobody kept a launch-day baseline to check.
A developer started one for Claude Opus 5.5 on release day. It's called livenerf, it reached the Hacker News front page overnight, and its validation run is the part worth reading.
Two identical runs differed by about 6 accuracy points. Dropping effort from high to medium moved accuracy about 4. The noise was bigger than the change.
Output tokens caught it. Medium effort cut them by about a quarter, low by about 60%. Swapping in the older Opus 5 wasn't distinguishable at all.
To catch drift in your own app:
Accuracy is noisy. Token counts are cheap evidence.