Anthropic / 7 October 2026

Anthropic Halves Sonnet 5.5’s Cached-Input Price

Anthropic cut Claude Sonnet 5.5’s cached-input price to US$0.10 per million tokens on 7 October, reducing the cost of repeatedly supplying material the model has already processed. The previous rate was US$0.20, while ordinary input and output prices remain US$2 and US$10 per million tokens respectively. The change targets a different part of application economics from a general model-price reduction: it benefits requests that successfully reuse cached context. For customers running agents against recurring instructions, reference documents or other stable material, the saving depends on how much of the bill comes from cache reads. Applications dominated by fresh input or generated output will receive a much smaller percentage reduction in their total spend.

At the new rate, reading 100 million cached tokens costs US$10 instead of US$20. That US$10 saving is exact for the stated quantity, but it says little about the complete job until the other charges are included. Suppose an illustrative workload also consumes one million ordinary input tokens and one million output tokens. Those components add US$12 at the published prices, leaving the combined bill at US$22 rather than US$32, a reduction of 31.25%. The example excludes the cost of creating the cache and any other services. It demonstrates why a 50% cut to one component need not become a 50% cut to the customer’s invoice.

Cache creation remains a separate economic consideration. Anthropic’s pricing documentation lists Sonnet 5.5 writes at US$2.50 per million tokens for a five-minute cache and US$4 for a one-hour cache. A customer pays that initial cost to make later reuse cheaper. In a simplified five-minute example, writing and then reading one million tokens once costs US$2.60, compared with US$4 for supplying that million-token quantity as ordinary input on two requests. The resulting US$1.40 difference assumes the material qualifies for reuse and is requested within the cache period. It is a calculation of the mechanism, not a claim that every agent achieves the same pattern.

Anthropic says the change makes Sonnet 5.5 run “around 20% cheaper on most agentic work.” That is the company’s workload claim, rather than a universal tariff rule or an independently established customer outcome. An application with no successful cache reads receives no saving from this particular cut. At the other extreme, frequent reuse of substantial context can make the reduction significant even when individual interactions are inexpensive. The important operating questions concern the proportion of reusable input, how often it changes and how reliably the application obtains cache hits. Those variables connect the headline price to actual expenditure and explain why two customers using the same model may see different results.

The change also affects the incentives for application design. Keeping stable instructions and reference material reusable can reduce the recurring cost of a long-running task; constantly reconstructing that material may sacrifice the benefit. Yet engineering work has its own price, and optimising a lightly used application can cost more than the tokens it saves. A practical investment comparison therefore puts expected request volume and cache reuse against the cost of maintaining the implementation. The price reduction improves that calculation for existing users without requiring a claim that the model has become more capable. It is a commercial adjustment to the way repeated computation is billed, with benefits concentrated in a recognisable workload category.

Analysis

Cached context is becoming a more important lever in competition for agent workloads. Anthropic can make repeated use cheaper while retaining the full output price, concentrating the concession on material that is economical to reuse. Customers with established, repetitive workflows gain most immediately; buyers evaluating occasional standalone questions gain little. The strategic effect depends on whether cheaper continuity encourages longer and more frequent tasks. If it does, additional output and fresh-input consumption can partly offset the discount, making the relevant measure revenue and service cost per completed workflow rather than the cache tariff in isolation.