Anthropic launches Claude Opus 5.5 – a flagship AI for everyone that’s cheaper and in many ways just as good as Fable 5.1

Anthropic launches Claude Opus 5.5 – a flagship AI for everyone that’s cheaper and in many ways just as good as Fable 5.1

Anthropic announces the Claude Opus 5.5 language model, the first in the new Claude 5.5 series. The company calls it its new flagship: the model is significantly better than the Claude Opus 5 in agent programming, computer work and application office tasks, while, according to the company, it costs 40% less to use on average.

    Image source: Human

Image source: Human

One of the main development directions for Opus 5.5 is complex and lengthy tasks where models must independently plan work, use tools, check intermediate results and make changes to large projects.

According to the developer, one early adopter used Claude Opus 5.5 to migrate a 680,000-line codebase in less than 24 hours—a manual task that he estimated would have taken a team of engineers several weeks. In another test, the model audited and patched a 200,000-line code base in less than three hours; the previous version of Claude Opus 5 required more than 20 hours and 2.5x the tokens to complete a similar job.

In a HAProxy migration test from C to Rust, Opus 5.5 completed the task in 9.5 hours, while Fable 5.1 took 12 hours, and the running cost was reduced by 51%. Anthropic also noted improvements in software optimization: In one test, Opus 5.5 was able to reduce load times for all pages of a web application in 39 out of 40 cases, while Opus 5 is more likely to make smaller changes that could alter the behavior of the application.

According to Anthropic’s internal testing, Claude Opus 5.5 scored 66.4% on the Terminal-Bench 4.0 benchmark, which evaluates an AI agent’s ability to perform multi-step tasks on the command line. This is higher than the results of Claude Opus 5, Fable 5.1, GPT-6 Astra and GPT-5.6 Sol. In FrontierCode v1.1, the new model scored 54.4%, while Claude Fable 5.1 scored 50.3% and GPT-6 Astra scored 53.3%. In the CursorBench 4.0 benchmark, Opus 5.5 scored 57.8%, better than Fable 5.1 and Opus 5.

Anthropic specifically notes that when it comes to flagship models, benchmark differences don’t always accurately reflect the actual differences between the models. By her own assessment, the actual gap between the Opus 5.5 and the Claude Fable 5.1 is smaller than the above test results suggest.

The developers also promise improved handling of text, research, financial models and business documents. In the GDPval-AA v2.1 test, which simulates professional tasks in 44 professions, Claude Opus 5.5 received a score of 1846 points. By comparison, Fable 5.1 has 1,735 and Opus 5 has 1,708.

In one of the internal tests, the model was asked to prepare a report on a company’s financial performance using a specially prepared copy of the network, where the required version was difficult to find. Opus 5.5 passed the company’s quality threshold (any errors in numbers or quotes will cause the test to fail) in 16 out of 18 attempts, while both Fable 5.1 and Opus 5 failed.

Anthropic has individually redesigned how the model communicates. Opus 5.5 improves response structure, capturing key messages immediately and reducing the use of jargon or unusual wording, the company said. This should make the model more convenient for long-term collaboration with users.

Meanwhile, Anthropic separately highlighted the effectiveness of the new model. The cost of one million input tokens dropped from $5 to $4, and the cost of output tokens dropped from $25 to $20. The cost of reading from the cache dropped from $0.50 to $0.20 per million coins, and the cost of writing went from $5 per million coins instead of $6.25. According to the company, fewer tokens are required to complete tasks, and the overall cost of using the model is reduced by an average of 40%. The power generation rate has also increased by more than 30%. Opus 5.5 also offers an accelerated fast mode: in Claude Code and Claude Platform it is 2.5 times faster than the standard speed, but costs $8 per million input tokens and $40 per million output tokens.

Prior to release, Claude Opus 5.5 underwent external testing, including Frontier Design and METR. Anthropic claims the Claude model is better than previous models at adhering to specified limits, and in tests trying to transcend isolation environments, it did so about 85 percent less frequently than Opus 5 and Mythos 5.1. At the same time, the company acknowledged that it is impossible to fully identify the risks of AI behavior before deployment.

Additional controls are required for sensitive requests in the areas of cybersecurity, biology, and model distillation. Most cybersecurity tasks will be redirected to Claude Opus 4.8, and certified experts will have expanded access to Opus 5.5 through the Cyber ​​Validation and Life Sciences Validation programs.

Opus 5.5 is said to be better resistant to instant injection attacks and less likely to take irreversible action. Additionally, Anthropic implements a mind-preserving mechanism designed to make it difficult to extract a model’s internal reasoning at scale through a network of fake API accounts.

Claude Opus 5.5 is now available in the Claude chatbot, API, and other Anthropic services. In the coming weeks, Anthropic also promises to release Claude Sonnet 5.5 and Claude Haiku 5.5, which will receive many of the same performance, efficiency, and security improvements.

If you find an error, select it with your mouse and press CTRL+ENTER.

Exit mobile version