Anthropic on Tuesday introduced Claude Opus 5.5, the first model in its new Claude 5.5 family, saying it performs at the level of Claude Fable 5.1 on most work while costing 40% less to run than Opus 5. The release was tested before launch by external evaluators, including Frontier Design and METR.
On the company’s automated behavioral audit, described as the most comprehensive alignment test Anthropic runs, Opus 5.5 is the strongest-performing model the company has tested to date. It also ships with the safeguards Anthropic has developed for its most capable models.
The performance claims are concrete. One early tester completed a 680,000-line code migration in less than a day, work Anthropic says would have taken an engineering team weeks. Asked to cut load times across every page of a web app, Opus 5.5 succeeded 39 of 40 times; Opus 5 made smaller improvements that also altered the app’s behavior. A different tester had several Claude models build a game from a single prompt, and Opus 5.5 scored higher than any other model on the strength of its graphics and polish.
Pricing reflects the reduced compute needed to serve the model. Input and output tokens are $4 and $20 per million, 20% less than Opus 5. Cache reads, which make up the majority of agentic and coding work costs, are $0.20 per million tokens, 60% less. Opus 5.5 also generates output more than 30% faster than Opus 5. Anthropic is raising five-hour usage limits on Pro, Max, Team and seat-based Enterprise plans, and giving subscription users a rate limit reset they can save and use whenever they choose. Fast mode, available in Claude Code and the Claude Platform, runs at up to 2.5x speed and costs $8 per million input tokens and $40 per million output tokens.
In an internal test, Anthropic asked Opus 5.5 and Fable 5.1 to translate HAProxy, widely used software that balances web traffic loads across servers, from C into Rust. Both rewrites passed nearly all of HAProxy’s own regression tests; Opus 5.5 finished in 9.5 hours against 12 for Fable 5.1, and cost 51% less. At its default effort level on FrontierCode, Opus 5.5 beats GPT-6 Astra at roughly 20% of the cost per task. On Terminal Bench 4.0 it matches Astra for about 40% of the cost, and on CursorBench it beats GPT-5.6 Sol by 11 points for about a third of the cost.
Safety testing shows similar movement. Opus 5.5 is much less likely than recent models to take hard-to-reverse actions or to act outside the boundaries it has been given, and is more resistant than Opus 5 to prompt injection, matching or beating it in every setting tested, including coding, tool use, computer use and web browsing. On a benchmark run by the security firm Gray Swan, it ties Fable 5.1 for the lowest prompt injection success rate of any model tested. In a new evaluation measuring a model’s propensity to cross containment boundaries, Opus 5.5 attempted to circumvent them around 85% less often than Opus 5 or Claude Mythos 5.1, and every attempt it made was low severity and self-reported. Anthropic also notes a limit: Opus 5.5 often suspects it is being evaluated, which the company says challenges its ability to assess how the model will act in real-world settings.
Because Opus 5.5 is comparable to Claude Mythos 5.1 in biology and cybersecurity, it is deployed with safeguards similar to those on Claude Fable 5.1. Most cybersecurity tasks will be re-routed to Opus 4.8, while vetted defenders can apply to an expanded Cyber Verification Program offering three tiers of increasingly permissive trusted access. Vetted organizations can apply to the Life Sciences Verification Program to use Opus 5.5 for biology research, work that includes a long-horizon molecular prediction and design evaluation conducted with Dyno Therapeutics. The model also launches with preserved thinking, the anti-distillation safeguard introduced with Fable 5.1, which stops API users from editing Claude’s prior context in an attempt to extract its reasoning; it applies to accounts created on or after August 31, 2026. Opus 5.5 is available with zero data retention and carries watermarking measures to comply with the EU AI Act. It is available on Amazon Web Services, Google Cloud and Microsoft Azure, and on the Claude Platform under the name claude-opus-5-5.
Early testers found the model’s writing clearer and easier to follow, feedback that addresses some of the common complaints about Opus 5. As one tester put it, “it writes the way I do.” Walleye Capital, an investment firm and early tester, reported that Opus 5.5 largely solved its evaluation suite on the lowest setting and, at higher settings, noticed an error in the firm’s evaluation instructions and corrected for it, something no other model had caught before.
On GDPval-AA v2.1, a test of real-world work across 44 occupations, Opus 5.5 scores 1846 Elo, ahead of Fable 5.1 and Opus 5, and at default effort it beats GPT-6 Astra at max effort for about a fifth of the cost per task. Anthropic said its own testing found the gap between Opus 5.5 and Fable 5.1 is narrower than benchmark margins suggest, and that efficiency is where the new model’s advantage is clearest.
Last week, Anthropic chief executive Dario Amodei argued that AI progress should be paced so that safety practices stay ahead of model capabilities. The company says its safety work now runs on two time horizons at once: established practices such as extensive alignment testing and pre-release evaluation for current models, and preparation for more advanced systems, including tighter filtering of reinforcement learning environments, improved alignment rewards and interpretability-based monitoring. Claude Sonnet 5.5 and Claude Haiku 5.5 will follow in the coming weeks, with many of the same improvements to performance, efficiency and safety.
