Opus 5.5 and GPT-6 Sol/Luna: the price war moved to cached tokens

Opus 5.5 cut cache reads 60% and GPT-6 gives 90% off cached input, so agent costs now hinge on a stable prompt prefix.

Opus 5.5 and GPT-6 Sol and Luna price changes with a checklist: cache hit rate, static context first, evals, thinking always on

Two frontier model launches landed within a day of each other. The headline is cheaper tokens. The real story is cached tokens.

Anthropic shipped Claude Opus 5.5. Input and output got 20% cheaper, but cache reads dropped 60%. Anthropic says cache reads make up most of the cost of agentic and coding work.

OpenAI shipped GPT-6 Sol and Luna at half the previous API price. It also gives 90% off cached input reads, and changing reasoning effort or tools no longer breaks the cache.

So an agent's bill now depends less on which model you pick and more on how stable your prompt prefix is.

What to check this week:

  1. Cache hit rate, per endpoint
  2. Static context first, user input last
  3. Your eval set, before swapping model ids
  4. Opus 5.5 can't run with thinking off. Test your integration.

Cheap tokens help. Reused tokens help more.

Watch the video on LinkedIn ↗