Anthropic Launches Haiku 5.5 With Cheaper Short Prompts
Anthropic released Claude Haiku 5.5 on 7 October, cutting the entry price for applications that need inexpensive, frequent model responses. Its direct API charges US$0.10 per million input tokens and US$0.50 per million output tokens for prompts of up to 100,000 tokens, one tenth of Haiku 4.5’s corresponding rates. Longer prompts attract higher prices, making the size of the material supplied to the model an important purchasing variable. The release extends Anthropic’s latest model generation into workloads where repeated small charges can determine whether an application is economical. Buyers gain a cheaper option for routine requests, while developers still have to establish whether it completes their particular tasks reliably enough to reduce total costs.
The pricing documentation separates the short and long categories: above the threshold, input costs US$0.50 and output US$2.50 per million tokens. That creates a fivefold rate difference, rather than a single universal discount. As a simple illustration, one million short requests using 1,000 input and 200 output tokens each would cost US$200 at the headline rates, before caching, retries or other charges. The same token volumes at Haiku 4.5’s listed input and output prices would cost US$2,000. This calculation holds token volumes fixed. Anthropic says the updated tokenizer uses slightly more tokens per task; response length and repeated attempts can also alter real spending.
Anthropic’s migration guide describes adaptive thinking as the default and an effort control that changes how much reasoning the model undertakes. Those controls give application designers another trade-off to evaluate alongside the token rate. A low-cost request that uses additional reasoning or requires repeated correction may be less attractive than a more expensive request that finishes successfully. The guide also warns that a small output allowance can be consumed by thinking before a visible answer appears. In operational terms, a model replacement therefore calls for testing the surrounding application’s limits and response handling, as well as checking the quality of the answer displayed to the end user.
Early customer evidence is encouraging but narrower than a general productivity claim. In Anthropic’s announcement, Asana’s Aaron Vinh described “a noticeably snappier experience,” alongside a reported latency reduction exceeding 30%. A faster response can matter in a product that waits for the model before displaying the next screen or completing an interaction. It does not, by itself, establish that the entire business process finishes 30% sooner. Waiting time may be only one component of the process, alongside data retrieval, human review and external systems. Customers evaluating the release can separate those components to determine whether a measured improvement affects their own service levels or merely one technical stage.
The model’s commercial role depends on the division of work within an application. Straightforward requests can be assigned to a cheaper model, with difficult cases escalated to a more capable and more costly alternative. Such routing only pays when the application can identify those cases without introducing excessive delay or misclassification. A useful comparison is the cost of an accepted result, including evaluation and correction, rather than the price of a token alone. Anthropic has supplied a substantially lower entry rate; the launch does not establish how many existing premium-model requests can safely move to Haiku, or how much additional demand a lower price will generate.
Analysis
Anthropic is exchanging revenue per token for the possibility of much greater request volume. At a 90% price reduction, an unchanged mix would require ten times as many billed tokens to preserve revenue, before any efficiency improvement or shift in workload. The stronger opportunity is therefore applications previously too expensive to operate at scale, rather than simply discounting existing traffic. Developers retain bargaining power if their routing layer can switch suppliers easily; Anthropic captures more value when Haiku’s reliability reduces the engineering and review costs that sit outside the API bill.