PostHog Jeeves: when reasoning pays off for a 9B classifier

Reasoning made Jeeves five points more accurate but pushed p90 latency to 17 s; reasoning only below 0.9 confidence keeps most of the gain.

Jeeves dev set: accuracy 0.775 to 0.825, median latency 0.3 s to 3.3 s and p90 17 s, with a checklist for classifiers

A 9B classifier got five points more accurate once it could think before answering. Its median response went from 0.3 seconds to 3.3.

PostHog open-sourced Jeeves today, a small decision model for yes/no, multiple-choice and rating questions. The kind of call behind ticket routing or content flags.

The tail is where it hurts. With full reasoning, p90 latency on one H100 reaches 17 seconds.

The middle setting is worth copying. The model answers without thinking first and only reasons, with a capped chain, when its confidence is below 0.9. On their dev set that keeps most of the gain, 0.806 against 0.825, at a 2 second median and under 6 at p90.

Before you add reasoning to a classifier:

  • Measure it on your own labels
  • Read p90, not the median
  • Skip thinking when the fast answer is confident
  • Check calibration before trusting the threshold

Reasoning is a budget. Spend it on the uncertain cases.

Watch the video on LinkedIn ↗