Franklin AI News Brief

Claude Haiku 5.5 adds adjustable effort as Anthropic cuts small-model costs

Key Takeaways

  • Anthropic reports lower average running costs for Haiku 5.5, introduces effort controls and halves Sonnet 5.5 cache-read pricing.
  • Benchmark results remain company-reported.
  • Anthropic has introduced Claude Haiku 5.5 for high-volume, cost-sensitive work, pairing a claimed reduction in running costs with a new adjustable effort setting.
  • The company positions the model for repetitive workloads, coding subagents and tasks where response speed matters, including customer support and browser use.
  • In its Haiku 5.5 announcement, Anthropic says the model costs around 75% less to run on average than Haiku 4.5.

Anthropic has introduced Claude Haiku 5.5 for high-volume, cost-sensitive work, pairing a claimed reduction in running costs with a new adjustable effort setting. The company positions the model for repetitive workloads, coding subagents and tasks where response speed matters, including customer support and browser use.

In its Haiku 5.5 announcement, Anthropic says the model costs around 75% less to run on average than Haiku 4.5. That is an average comparison from the company, not a guarantee that every application will see the same saving. The captured announcement does not provide a complete token-price table or the footnote methodology needed to reproduce that percentage independently.

Effort becomes a control on the smallest Claude tier

Haiku 5.5 is Anthropic's first Haiku-class model with an adjustable effort setting. Users can choose how to balance cost and intelligence, rather than treating every request as requiring the same amount of work. The announcement illustrates that trade-off using computer-use, knowledge-work and reasoning evaluations.

The proposed workloads include summaries, compactions, database queries and classification requests. Anthropic also describes Haiku working alongside Opus 5.5 and Sonnet 5.5 as a coding subagent. Those are product positioning claims, not evidence that the model will correctly complete every database operation or software task without supervision.

For an application processing many small requests, the useful question is which effort level meets its quality requirement. A cheaper attempt that frequently needs a retry can produce a different total cost from the headline comparison. Evaluating complete task outcomes, latency and retries together is more informative than selecting a model solely from its average saving.

The benchmark labels matter as much as the numbers

Anthropic reports a 72.4% result on the OSWorld 2.1 offline subset for Haiku 5.5, compared with 15.7% for Haiku 4.5 in the same displayed table. The offline-subset qualification should travel with the number. It is not a universal success rate for using arbitrary computers or websites.

The table also lists 39.2% on Terminal-Bench 4.0 and 46.4% on FrontierCode 1.1 Main. These are different coding evaluations, not interchangeable measurements of the same workload. On Humanity's Last Exam, the displayed results separate a 45.9% no-tools score from 57.4% with tools. Combining those settings into one unspecified reasoning score would remove an important condition of the evaluation.

All of these figures come from Anthropic's release material. They offer starting points for choosing what to test, but Franklin has not reproduced them. A team's own evaluation should use the tasks and failure costs that actually matter in its deployment.

Sonnet cache reads and subscriber credits change too

Alongside the Haiku release, Anthropic is halving the price of Claude Sonnet 5.5's cache reads. The company says that makes Sonnet around 20% cheaper on most agentic work. That estimate depends on the workload; applications with a different amount of reusable context should not assume the same reduction.

The earlier Sonnet 5.5 launch focused on faster output and lower task costs. This announcement adds a distinct cache-read price change, so those two claims should not be collapsed into a single explanation of how the savings occur. Anthropic also announces monthly API credits for Claude Max and Team subscribers, without a complete credit-amount table in the captured material.

Asana supplies one customer example. Staff software engineer Aaron Vinh says its AI Teammates evaluation covered bug triage, project setup and portfolio searches, reporting over a 30% reduction in task-completion latency and up to 2.5-times faster inference per agent turn compared with its current model. That is a customer-reported result from a particular evaluation suite, not an independently verified promise for other applications.

Our read

Franklin AI Take

Adjustable effort is the most useful new control here: Haiku 5.5 can trade cost against intelligence instead of applying one setting to every request. Anthropic also says it is halving Sonnet 5.5 cache-read pricing. Test those changes on complete tasks, including retries and incorrect outputs. The company's average savings and benchmark scores do not establish the cost or reliability of your own workflow.