Cursor cut its agent system prompt by two thirds; token cost fell 7%

Most of Cursor's agent spend sits in growing context, not the cached system prompt; on-demand tools and cache breakpoints moved the bill.

Cursor agent harness, 7% fewer tokens: system prompt trimmed by about two thirds, and the four changes that moved the bill

Cursor cut about two thirds of its agent's system prompt. Total token cost fell 7%.

That gap is the useful part.

The system prompt is the part a team fully controls, so it gets trimmed first. It's also resent every turn and mostly served from cache. The rest of the bill sits in what piles up: file reads, tool output, a conversation that keeps growing.

What moved their number:

  • Tools loaded on demand. Most built-in tools show up in under a fifth of conversations, so their definitions left the static context.
  • Cache breakpoints after the stable layers, with per-request setup moved behind them. Cold cache misses fell by a fifth.
  • Line numbers on every tenth line of a file read.
  • A/B tests on real traffic, since evals skew toward hard tasks.

If you run your own agent harness, split spend by layer before you trim anything.

Trim the prompt for hygiene. Trim the context for money.

Watch the video on LinkedIn ↗