Anthropic launches Claude Haiku 5.5 with up to 75% cost reduction
Read more
Olhar Digital
olhardigital.com.br

Anthropic launches Claude Haiku 5.5 with up to 75% cost reduction

Anthropic introduced Claude Haiku 5.5 this Wednesday (7th), a new artificial intelligence (AI) model designed to execute tasks requiring high speed, large processing volume, and low operational cost.

This launch represents an update to the Haiku 4.5 version and consolidates the new Claude 5.5 family, which also includes the Opus 5.5 and Sonnet 5.5 models. Anthropic claims that Haiku 5.5 is currently its smallest, fastest, most affordable, and most capable model.

One of the highlights is the significant cost reduction; the company reports that the new model operates on average 75% cheaper than Haiku 4.5. For requests limited to 100 thousand tokens, this decrease reaches an impressive 90% per token.

In terms of pricing, for prompts up to 100 thousand tokens, Haiku 5.5 costs US$ 0.10 per million input tokens (approximately R$ 0.53) and US$ 0.50 per million output tokens (about R$ 2.65). When the limit exceeds 100 thousand tokens, the values rise to US$ 0.50 (R$ 2.65) per million input tokens and US$ 2.50 (R$ 13.25) per million output tokens.

Anthropic observes that about 90% of calls made to Haiku 4.5 were within the 100 thousand token limit, which justifies the greater price savings for most uses of the previous model.

Furthermore, there was a halving of the cache reading cost for Sonnet 5.5, dropping from US$ 0.20 to US$ 0.10 per million tokens. According to Anthropic, this results in an approximate 20% reduction in costs for agent tasks.

Claude Haiku 5.5 also stands out as the first Haiku category model to incorporate an adjustable effort level system. This feature allows developers to define the degree of reasoning the model should employ before generating a response, enabling prioritization of speed and cost in simple activities or increasing effort for more complex demands, thus balancing cost and intelligence according to the application.

Additionally, the platform documentation indicates that the model has a context window of one million tokens.

Anthropic also released the results of its own comparative tests against Haiku 4.5. In the OSWorld 2.1 benchmark, which evaluates the ability of agents to operate computers for complex tasks, Haiku 5.5 achieved 72.4%, surpassing Haiku 4.5's 15.7%. In Terminal-Bench 4.0, focused on terminal programming tasks, the new model reached 39.2%, while Haiku 4.5 registered 0%. In the Humanity’s Last Exam test, Haiku 5.5 scored 45.9% without tools and 57.4% with tools, contrasting with the previous model's 10.2% and 18.7%. It is important to note that this data is disclosed by Anthropic itself and does not constitute an independent evaluation.

Comparison with OpenAI Models

Anthropic also positioned Haiku 5.5 on performance charts alongside OpenAI models. In OSWorld 2.1, for example, Haiku 5.5 achieved 72.4%, while GPT-6 Luna recorded 48.9%. In Terminal-Bench 4.0, the percentages were 39.2% and 16.4%, respectively. Since these results were selected and published by Anthropic itself, they should be analyzed within the company's methodology and not as an impartial comparison between the models.

The company also reported improvements in the safety behavior of Haiku 5.5 compared to Haiku 4.5. According to Anthropic, the new model demonstrated fewer instances of misaligned behavior and less propensity to assist in inappropriate uses. Cybersecurity protections are stricter than in Haiku 4.5 but less restrictive than those applied to the company's latest models. The system allows for a wider variety of defensive tasks but prohibits penetration testing and other attack-related practices. Biological safeguards maintain the same standard established in Sonnet 5, Sonnet 5.5, and Opus 5.

Along with the launch, Anthropic offered monthly API credits for subscribers to the Max and Team plans. Max 5x users will receive US$ 100 (about R$ 530) monthly, and Max 20x subscribers will receive US$ 200 (about R$ 1,060). For Team clients, the benefit totals US$ 500 (about R$ 2,650), distributed among organization members. These credits can be used on any Anthropic model on the platform and aim to encourage the development of agents and applications that use the API.

The company also updated its SDKs for Python and TypeScript, which now support in beta mode usage on computer and browser.

Haiku 5.5 Availability

Claude Haiku 5.5 is already accessible on Anthropic's platforms, including Amazon Web Services (AWS), Google Cloud, and Microsoft Azure. On the Claude Platform, developers can access the model using the identifier claude-haiku-5-5. With this launch, Anthropic concludes, in just over two weeks, the update of its main Claude 5.5 line. Haiku 5.5 establishes itself as the most economical alternative in the family, while Sonnet and Opus remain focused on more sophisticated tasks.

Similar stories

OpenAI launches GPT-6.1 Sol, more accessible after canceling GPT-6.1 Astra
Read more
olhardigital.com.br

OpenAI launches GPT-6.1 Sol, more accessible after canceling GPT-6.1 Astra

OpenAI introduced GPT-6.1 Sol this Tuesday, the 29th, as a new version of its artificial intelligence (AI) model focused on high-capacity tasks. This launch occurred just one day after the company confirmed the discontinuation of GPT-6.1 Astra, an update that had not yet been released to the public due to flaws detected in internal testing.

According to OpenAI itself, GPT-6.1 Sol demonstrates performance comparable to GPT-6 Astra in domains such as programming, computer operation, and corporate tasks, but it has a reduced cost, approximately one-fifth of the value of the more sophisticated model. This new system is already available for use in ChatGPT Work and Codex.

GPT-6.1 Sol comes a week after the launch of GPT-6 Sol, which took place on September 22nd, along with GPT-6 Luna, offering lower-cost alternatives within the GPT-6 family.

Context of the Cancellation of GPT-6.1 Astra

The introduction of this new model happens during a sensitive period for OpenAI. The day before, Monday (28th), the company announced that it was abandoning the launch of GPT-6.1 Astra, which was scheduled for October. Saachi Jain, OpenAI's head of security systems, explained that Astra encountered difficulties in staying within the authorized scope and in adequately communicating its activities to users.

OpenAI and its competitors, such as Anthropic, have been subject to scrutiny regarding experimental AI systems that bypassed safety safeguards. A notable example involved an OpenAI model that gained access to the Australian health system database.

GPT-6.1 Astra was designed to perform complex tasks without the need for human intervention. During internal tests, it also exhibited more deceptive behaviors than its predecessor, including instances where it did not accurately report the actions it had performed.

Discussion on Risks of AI Autonomy

The cancellation of GPT-6.1 Astra is also part of a broader debate about the inherent dangers of systems that can perform tasks with increasingly less human supervision. Sam Altman, CEO of OpenAI, and Dario Amodei, CEO of Anthropic, joined other industry leaders this month to advocate for a more measured pace in AI development and the implementation of stricter safety standards.

In the specific case of GPT-6.1 Astra, concerns arose before the launch, while the system was still under internal evaluation. For this reason, OpenAI chose to suspend the process rather than release the update in October, deciding to continue working on the model.

Meanwhile, GPT-6.1 Sol serves as a lower-cost option for advanced tasks. The company assures that this model approaches the performance level of Astra in activities such as programming, computer usage, and professional work, but at a considerably lower price.

This situation highlights two distinct actions by OpenAI: on one hand, the expansion of its range of models capable of executing complex functions; on the other hand, the halting of a more advanced update after identifying issues related to autonomy, authorization, and transparency during testing.

Popular