Anthropic has launched Claude Opus 5.5, the first model in its new Claude 5.5 generation, with a clear bet on agents capable of carrying out long stretches of coding, research, and computer-use work. The company says it beats Opus 5 on several benchmarks and cuts the cost of common workloads by 40%, while adding new safeguards to control autonomous actions and high-risk tasks.
Claude Opus 5.5 in 20 seconds
- Claude Opus 5.5 launches the Claude 5.5 family and is available now.
- Anthropic puts the savings versus Opus 5 at 40% on common workloads.
- It costs $4 per million input tokens and $20 per million output tokens.
- It improves on agentic coding, professional work, and computer use.
- Sonnet 5.5 and Haiku 5.5 are coming in the following weeks.
The move comes at a moment when leading models are no longer competing solely on answering a single question well. Attention is shifting toward systems that can sustain a task for hours, use tools, edit code, review results, and complete processes with less human intervention. Anthropic has built Opus 5.5 around that scenario.
The model is available at launch on Claude, Claude Platform, and the major cloud services that distribute Anthropic’s models: Amazon Web Services, Google Cloud, and Microsoft Azure. On the developer platform, it’s identified as claude-opus-5-5.
Lower cost for token-heavy work
One of the figures Anthropic emphasizes most isn’t in the benchmarks, but on the bill. Opus 5.5 is priced at $4 per million input tokens and $20 per million output tokens. Opus 5 cost $5 and $25, respectively.
The difference is even bigger for cache reads, which are common in agentic and coding workflows. Opus 5.5 charges $0.20 per million tokens versus $0.50 for Opus 5. Cache writes drop from $6.25 to $5.
Anthropic calculates that these changes, combined with using fewer tokens to complete certain tasks, cut execution costs by roughly 40% compared with Opus 5 on common workloads.
The company also says Opus 5.5 generates responses more than 30% faster than its predecessor. For developers running agents over long stretches, cost per task can matter more than the price of any single token.
There’s also a fast mode for Claude Code and Claude Platform that can reach up to 2.5 times the speed of standard mode. It’s priced at $8 per million input tokens and $40 for output.
The focus is on autonomous coding
Coding is one of the areas where Anthropic puts its strongest arguments. On Terminal-Bench 4.0, a benchmark focused on complex, professional command-line tasks, Opus 5.5 scores 66.4%, versus 55.8% for Fable 5.1 and 52.3% for Opus 5.
It reaches 54.4% on FrontierCode v1.1 and 57.8% on CursorBench 4.0. Anthropic also compares its results against models from other vendors, though it cautions that differences between the most advanced models don’t always carry over directly into the real-world experience of using them.
There are examples closer to a development team’s everyday work. According to Anthropic, an evaluator used Opus 5.5 to review and fix a 200,000-line codebase in under three hours. Opus 5 took more than 20 hours on the same test and used 2.5 times more tokens.
In another test, the model tackled a 680,000-line code migration that, per the evaluator’s estimate cited by Anthropic, would have taken an engineering team weeks of work.
The company also tested Opus 5.5 with HAProxy, a C software project used to balance traffic across servers. The goal was to translate it to Rust. Both Opus 5.5 and Fable 5.1 passed nearly all of HAProxy’s regression tests, but the new model finished the job in 9.5 hours versus 12 hours for Fable 5.1, and at 51% lower cost, according to Anthropic’s internal test.
These figures come from the company’s own evaluations or early-access testing. They aren’t equivalent to an independent measurement across every development environment, something Anthropic itself also acknowledges when it flags the limits of benchmarks once models reach high levels of capability.
It also targets office work
Opus 5.5 isn’t built only for writing code. Anthropic presents it as a model for professional tasks that combine research, analysis, and generating documents, spreadsheets, and presentations.
On GDPval-AA v2.1, which evaluates professional work across 44 occupations, Opus 5.5 scored 1,846 Elo points under the maximum-capability conditions Anthropic used. Fable 5.1 scored 1,735, and Opus 5 scored 1,708.
The company also ran a test tied to financial analysis. Opus 5.5 and Opus 5 were tasked with evaluating a possible acquisition between two fictional HR-software companies. Both models reached similar conclusions, but Anthropic says the new model produced a more complete financial model and a presentation that was easier to review.
Opus 5.5 took 63 minutes versus 93 minutes for Opus 5 and, according to the company, produced the result at 50% lower cost.
Another change affects communication. Anthropic says it specifically worked to have the model lead with the relevant information, use less jargon, and better follow the writing rules a user provides. It’s a less visible improvement on a benchmark table, but a particularly meaningful one in long agent sessions that generate large amounts of information.
Safety gains weight as autonomy grows
Anthropic paired the launch with a battery of behavioral evaluations. The company says Opus 5.5 achieved its best results to date on its automated behavioral audit, which examines close to 2,000 scenarios.
In a test designed to measure attempts to bypass containment limits, Anthropic says Opus 5.5 attempted to get around those barriers roughly 85% less often than Opus 5 and Claude Mythos 5.1. The company also notes that the attempts it detected were low-severity and that the model reported them.
The company does acknowledge one important limitation: a model can detect that it’s being evaluated. Because of that, Anthropic considers that alignment tests alone don’t provide a guarantee of behavior across every real-world environment.
The model includes specific safeguards for cybersecurity and biology. For cybersecurity tasks, Anthropic says a large share of work will be routed to Opus 4.8, while verified professionals will gain progressive access to broader capabilities through its verification program.
For biology, Opus 5.5 uses measures similar to those in Fable 5.1. Verified organizations can request access through the Life Sciences Verification Program.
It also includes measures against distillation attacks, in which a third party tries to extract a model’s capabilities using large numbers of accounts. Opus 5.5 uses the preserved-thinking system Anthropic introduced previously to prevent certain API users from manipulating Claude’s prior context for that purpose.
The company also maintains zero-data-retention options and has added labeling measures tied to compliance with the European Union’s AI Act.
Claude Opus 5.5 therefore arrives with a proposition that combines two variables that usually compete with each other: more capacity for long-running tasks and lower operating cost. Anthropic has also confirmed the family will continue with Claude Sonnet 5.5 and Claude Haiku 5.5, expected in the coming weeks.
FAQ
What does Claude Opus 5.5 improve over Opus 5?
Anthropic highlights improvements in agentic coding, professional work, computer use, speed, and communication, along with lower token consumption on certain tasks.
How much does Claude Opus 5.5 cost?
The standard price is $4 per million input tokens and $20 per million output tokens. Cache reads cost $0.20 per million.
Where can Claude Opus 5.5 be used?
It’s available on Anthropic’s own platforms and through Amazon Web Services, Google Cloud, and Microsoft Azure.
Will there be more Claude 5.5 models?
Yes. Anthropic has announced Claude Sonnet 5.5 and Claude Haiku 5.5 for the coming weeks.

