Anthropic launches the Claude Sonnet 5.5, the second model in the new Claude 5.5 series following the flagship Claude Opus 5.5. Developers are positioning the new product as a faster, more economical model suitable for everyday tasks, programming, processing documents and creating presentations.
Image source: Human
The main change compared to Claude Sonnet 5 is a significant increase in performance while reducing the number of calculations required to perform common tasks. According to Anthropic, Sonnet 5.5 generates responses more than 30% faster than previous versions and is up to 30% cheaper in most cases—the new model typically uses fewer tokens to perform the same task. At the same time, the price of a token does not change: $2 for 1 million input tokens and $10 for 1 million output tokens. The cost of reading tokens from the cache is $0.20 per 1 million tokens. By the way, Claude Sonnet 5.5 supports contextual windows up to 1 million tokens and withdrawals up to 128,000 tokens.
As for performance metrics, Anthropic claims growth in programming has been particularly significant. The Terminal-Bench 4.0 test evaluates the ability of artificial intelligence to independently perform multi-step tasks on the command line. Sonnet 5.5 scored 70.6%, while Sonnet 5 scored only 10.3%. In CursorBench 4.0, the new model scored 55.5%, while its predecessor scored only 34.1%.
In FrontierCode 1.1, the new Sonnet 5.5 scored 46.2% on maximum effort, while the Sonnet 5 scored 42.4%. The flagship Claude Opus 5.5 scored 54.4% in this test. Anthropic notes that in some tests, at best effort, Sonnet 5.5 has come close to the capabilities of the more expensive Opus 5.5, although the latter retains its advantage in complex open work that requires long periods of reasoning and independent decision-making. Let us remind you that Opus 5.5 costs twice as much: $4 and $20 per million input and output tokens respectively.
Sonnet 5.5 isn’t just for developers, however. Anthropic emphasizes processing documents, presentations, and spreadsheets, as well as building user interfaces. The company says the new model better follows the demo template and is able to produce results that require less manual post-processing.
Working on long tasks has also been significantly improved. The GDPval-AA v2.1 test simulated real professional work performance in 44 professions and 9 industries, with Sonnet 5.5 scoring 1844 points and Sonnet 5 scoring 1449 points. The result is very close to Opus 5.5 – 1846 points. In AA-Briefcase testing, the new model was also significantly ahead of its predecessor: 1811 points versus 1359 points.
Anthropic singled out improvements in imaging and computer interface processing. In the OSWorld 2.1 test, which evaluates the ability of artificial intelligence to interact with operating systems, Sonnet 5.5 scored 80.1%, while Sonnet 5 scored 57%. The model also got better at recognizing graphics and visual information. Sonnet 5.5 is the first mod in the Sonnet series to be able to complete the game Pokémon Red using only screenshots as visual information.
The changes also affect security systems. Thanks to significant enhancements in network security capabilities, Sonnet 5.5 is the first model in the Sonnet series to feature protection mechanisms and fallback limits similar to those used by the more powerful Anthropic models at launch. For high-risk cybersecurity tasks, the system can switch to Sonnet 5.
In addition, Sonnet 5.5 also gains protection against distillation – the attempt to use batch requests from a large number of accounts to extract model data and then use it to train another high-performance model. Anthropic uses a special classifier in the new version that prevents the extraction of internal reasoning.
Claude Sonnet 5.5 is now available on the Claude chatbot and on all major platforms, including Claude’s own platform, Amazon Web Services, Google Cloud and Microsoft Azure. Developers can use the model through an API.
In the coming weeks, Anthropic is also promising the Claude Haiku 5.5, another model in the Claude 5.5 series aimed at high-load and cost-sensitive scenarios.
If you find an error, select it with your mouse and press CTRL+ENTER.










