
Claude Opus 5, Anthropic's new model that nearly matches the flagship at half the price
Anthropic released Claude Opus 5, a model that comes close to the flagship Fable 5 in intelligence while costing half as much. Let's break down where it beats the previous version, why it's called the company's most reliable model yet, and what limits it still has.
What Claude Opus 5 is and who it's for
Anthropic has opened access to Claude Opus 5. The company describes the model as thoughtful and proactive, meaning it doesn't just hand over the first answer that comes up, it verifies its own work and carries the task through to the finish. In terms of intelligence, Opus 5 comes very close to the flagship Claude Fable 5 while costing half as much.
On coding and knowledge work benchmarks like Frontier-Bench and GDPval-AA, Opus 5 posts the best results of any model on the market. The one area where it falls behind is cybersecurity, where the more specialized Mythos 5 model still leads.
The core idea behind this launch is that the model was built for everyday use rather than occasional heavy tasks. It works more efficiently than other models, which is why it became the default model on the Claude Max plan and the strongest model available on Claude Pro.

Why Opus 5 is a better deal than the previous Opus 4.8
Claude Opus 5 delivers noticeably stronger results at the same price as the previous Opus 4.8. The model has an effort setting that lets you choose what matters more at any given moment, maximum intelligence or saving tokens for a faster and cheaper answer.
The model performs especially well on software development tasks. On the Frontier-Bench v0.1 benchmark, Opus 5 beats every other model and more than doubles Opus 4.8's result, with each task costing less than before. On CursorBench 3.2 at max effort, the model lands within 0.5 percent of Fable 5's best score while costing half as much per task.

These results are from an internal run of Frontier-Bench v0.1, on the mini-SWE-agent harness and a GKE backend, mean reward over 5 attempts per task. Opus 4.8 served as fallback on safety-classifier refusals for Opus 5 and Fable 5.
So the benefit here works twice over. The model both handles the task better and does it while burning fewer resources for the same result.
Where Claude Opus 5 beats every other model
The same pattern shows up in knowledge work and in solving unfamiliar problems. Here are the most notable results,
On ARC-AGI 3, a test where the model has to solve problems it has never seen before, Opus 5 scores three times higher than the next-best model
On Zapier AutomationBench, which checks whether a model can carry a real business task from start to finish, Opus 5 succeeds roughly 1.5 times more often than the next-best model at the same cost per task
Even at its lowest effort setting, Opus 5 passes more tasks than any other model
On the OSWorld 2.0 computer-use benchmark, the model outperforms everything else at any given price, surpassing Fable 5's best result at just over a third of the cost




Effort ladders, unless otherwise noted (low, medium, high, xhigh, max)
In plain terms, Opus 5 wins not only on raw quality but on the ratio between quality and price. Even when a competitor posts a similar score, Opus 5 usually gets there for noticeably less money.
How Opus 5 handles scientific tasks and visual output
In scientific research, Opus 5 is a real step up from Opus 4.8. The model posts better results on every one of the life sciences evaluations, covering structural biology, organic chemistry, and bioinformatics.
The biggest gains show up in organic chemistry, for example in tasks where the model has to work out a molecule's structure from spectroscopy data, where it scores 10.2 percentage points higher than Opus 4.8. On protein-related tasks, such as predicting how changes in a protein's sequence affect the way it functions, the gain is 7.7 percentage points. Worth calling out separately, the model's visual output got noticeably stronger too.
Why Opus 5 sees tasks through, three real examples
The key difference with Opus 5 is that it verifies its own work far more carefully and keeps trying until it actually succeeds. Here are three examples from testing and early access that show this well.
In the first case, the model was given a drawing of a machine part and asked to write code that would rebuild it as a 3D model. The catch was that it had no way to look at the drawing directly, that ability was deliberately removed. Opus 5 responded by writing its own computer vision pipeline to pull the geometry straight out of the raw pixels, then reconstructed the full part. It repeated that success several times over, while no competing model with the same setup could solve it even after five attempts.
In the second case, the model was handed a real bug in a popular package manager. Opus 5 found the root cause and fixed an edge case that the developer community's own patch had missed. A competing model fixed only the surface symptom without touching the underlying cause, then reported the bug as resolved.
The third example comes from an engineer at a trading firm who built a market data feed for a new exchange in a single session using Opus 5. Finding no live data stream to validate against, the model built its own test harness to confirm that its code parsed the exchange's data correctly.
How safe and predictable Opus 5 is
Before release, Anthropic ran an automated behavioral audit, and the results showed Opus 5 to be the company's most aligned model to date. In plain terms, it follows the rules built into Claude's constitution better than the others, shows the lowest rate of deceptive behavior, and is the hardest to trick into doing something it shouldn't.
The model also became the most careful when it comes to risky actions with hard-to-reverse consequences. On the overall misaligned behavior score, Opus 5 comes in at 2.3, the lowest and therefore best result among the company's recent models.

On our automated behavioral audit, Opus 5 scores 2.3 on overall misaligned behavior, the lowest of our recent models
What limits Opus 5 has in cybersecurity
Anthropic makes a point of noting that Opus 5 does not push the frontier on dual-use capabilities that could be misused. In evaluations run alongside private-sector and government partners, the model stayed behind Mythos 5 in both biology research and offensive cybersecurity.
An interesting detail here, the model was deliberately not trained on cyber tasks, yet it still improved substantially in that area simply because it got more capable overall. At finding vulnerabilities, Opus 5 comes close to Mythos 5, but at turning a discovered vulnerability into a working attack, it falls significantly behind.

On OSS-Fuzz, one of our cybersecurity evaluations, Opus 5 is close to Mythos 5 at identifying software vulnerabilities (left), but is considerably less successful at developing exploits for them (right)
Opus 5's safeguards are built so they don't get in the way of legitimate work. The model is allowed to look for vulnerabilities in source code, while binary-based vulnerability scanning, penetration testing, and exploit generation are blocked. By the company's estimate, these restrictions should trigger roughly 85 percent less often than on Fable 5. On the biology side, Opus 5 is now the company's most capable generally available model for scientific research, though it still has real limits on long-running, autonomous research tasks.
How much Claude Opus 5 costs and where to try it
Claude Opus 5 is available today across all platforms. It costs 5 dollars per million input tokens and 25 dollars per million output tokens, exactly the same as the previous Opus 4.8. Developers can start using it through the Claude API.
There's also a Fast mode, where the model runs roughly 2.5 times faster than default. It's available at twice the model's base price, and in Claude Code it can be paid for with usage credits.
Alongside Opus 5, the company shipped two updates in beta. The first lets developers change which tools the model can use in the middle of a conversation without invalidating the prompt cache. The second allows requests flagged by the safety system to be automatically routed to a different model instead of simply being blocked.
Smart, Careful, Affordable
Claude Opus 5 reflects the same trend as other recent releases on the market, the competition is no longer just about the highest benchmark score, it's about who delivers comparable results for noticeably less. The model came close to the flagship Fable 5 in intelligence at half the price, became the company's most careful and predictable model, and verifies its own work well enough to carry tasks through to the finish on its own.
That said, a perfect all-purpose model still doesn't exist. Mythos 5 remains ahead in cybersecurity, and for certain tasks models from other companies will work out better. There's only one way to find out, run your own real task across different models and compare the results rather than relying on someone else's rankings.
The problem is, setting up a separate subscription with every company just to run that comparison takes too much time and money. It's far more practical when every current AI model sits in one place and switching between them takes a couple of clicks.
That's exactly what unitool.ai is for. Get one subscription and unlock access to the newest models at once, including Anthropic's latest releases, no foreign cards and no separate sign-up on each platform. See for yourself which model turns out more accurate and more cost-effective for your specific tasks.