livenerf on Claude Opus 5.5: output tokens catch model drift first

livenerf's validation found accuracy noise bigger than the real change, while output token counts clearly showed effort shifts.

livenerf validation on Opus 5.5: medium vs high effort cut accuracy 4 points and tokens 26%, while identical runs were 6 points apart

Every few weeks someone swears a model got worse after launch. Almost nobody kept a launch-day baseline to check.

A developer started one for Claude Opus 5.5 on release day. It's called livenerf, it reached the Hacker News front page overnight, and its validation run is the part worth reading.

Two identical runs differed by about 6 accuracy points. Dropping effort from high to medium moved accuracy about 4. The noise was bigger than the change.

Output tokens caught it. Medium effort cut them by about a quarter, low by about 60%. Swapping in the older Opus 5 wasn't distinguishable at all.

To catch drift in your own app:

  1. Log output tokens per request next to eval scores.
  2. Pin your harness and CLI version. A changed harness looks like a changed model.
  3. Keep a frozen set of questions the model only sometimes gets right.

Accuracy is noisy. Token counts are cheap evidence.

Watch the video on LinkedIn ↗