Anthropic’s Opus 5 cuts coding cost in half, priced at $5 per million input tokens
A cheaper Claude model aims to win the enterprise middle ground, where token efficiency beats raw frontier power.

Anthropic launched Claude Opus 5 on Friday, positioning it as nearly as capable as its top model Claude Fable 5 while costing half as much. For enterprise buyers, the shift is about inference economics, verification behavior, and reducing the hidden cost of human review.
Anthropic launched Claude Opus 5 on Friday, and it comes with a blunt pricing signal: $5 per million input tokens and $25 per million output tokens, unchanged from its predecessor, Opus 4.8. The company says Opus 5 delivers nearly all the intelligence of its top-of-the-line Claude Fable 5, at half the cost, and it is being rolled out immediately across Anthropic’s platforms.
This is not a “we made the smartest model” press release. Anthropic is explicitly using Opus 5 as the default model for Claude Max and the strongest model available on Claude Pro. The real claim is operational: that most of the useful AI work enterprises run lives in a middle band of difficulty, where near-frontier intelligence delivered efficiently and cheaply beats frontier intelligence delivered expensively. Anthropic’s own spokesperson framed it as a lineup with jobs attached. Opus 5 is for “your daily driver,” Fable 5 is for the longest, most ambitious autonomous projects, Sonnet 5 is for scaled work where speed and cost per call decide what ships, and Haiku 4.5 is for subagents and instant answers.
So how does Opus 5 actually stack up? On paper, the numbers look like a direct challenge to the idea that buyers should always pay for the absolute highest capability. Anthropic says Opus 5 sets new state-of-the-art marks on coding and knowledge-work evaluations including Frontier-Bench and GDPval-AA. On Frontier-Bench v0.1, an agentic terminal coding benchmark, Opus 5 scores 43.3 percent, more than double Opus 4.8’s 18.7 percent, and ahead of Fable 5’s 33.7 percent, while also coming at a lower cost per task, according to the company. On ARC-AGI 3, an evaluation of novel problem-solving, Anthropic reports Opus 5 scored three times as high as the next best model. On OSWorld 2.0, a computer-use benchmark, Anthropic says Opus 5 surpasses Fable 5’s best result at just over a third of the cost.
But Anthropic also includes caveats, and in an industry that often treats benchmarks like gospel, those caveats matter. Opus 5 remains behind Mythos 5 on cybersecurity tasks and biology research, and an OpenAI-family model still leads on one agentic coding benchmark. More importantly, when asked where Opus 5 still falls short of Fable 5, the spokesperson’s answer amounted to a candid admission about what benchmarks do not measure. The evals where Opus 5 wins are bounded tasks with a specific outcome, which is where it is strongest. What they do not measure is duration. In plain English, Opus 5 is best when the job fits inside the evaluation box. Fable 5 is the one you reach for when the job outruns the benchmark.
That bounded-task versus long-horizon autonomy framing may become a major sorting mechanism as 2026 approaches and benchmark coverage saturates. Anthropic’s spokesperson advised customers to “run both on a representative workload,” meaning one bounded task and one long-horizon job. The reason is simple: a model that aces short puzzles can still struggle when it has to stay coherent across many connected steps over hours or days, with dense source material. For buyers, that is a procurement question, not just a curiosity. It changes how you design workflows, how you allocate budget, and how you decide which tasks are safe to automate end-to-end.
Even the way Opus 5 is marketed reflects that enterprise reality. Anthropic threads a theme through the launch: token efficiency is the real battleground for AI spending. Enterprises do not just pay for “capability.” They pay to run the model at scale. Opus 5 includes an adjustable “effort” setting so customers can trade intelligence for speed and token savings. Anthropic’s charts emphasize performance at a given cost rather than peak performance alone.
The launch also leans into another enterprise pain point that often hides behind raw scores: verification. Anthropic says Opus 5 verifies its work and iterates until it succeeds, and it provided examples meant to show behavior, not just outputs. In one Frontier-Bench task, the model was asked to reconstruct a machine part as a 3D CAD model from a drawing it was given no way to view. Instead of stalling or failing, Anthropic says Opus 5 wrote its own computer vision pipeline to extract geometry from raw pixels, repeatedly, while no competing model solved the task in five attempts. In another case, Anthropic says Opus 5 found the root cause of a real bug in a popular open-source package manager and fixed an edge case that the community’s own patch missed, while a competing model patched only the symptom.
Customers described similar “it checks itself” behavior. Harvey, a legal AI company, said Opus 5 achieved similar performance to Opus 4.8’s maximum-reasoning mode while generating 26 percent fewer tokens on average, according to Niko Grupen, head of applied research at Harvey. Richard Pham of Fundamental Research Lab said that on hard financial-modeling tasks, Opus 5 averaged nine percentage points higher accuracy while using roughly one-third fewer turns and tool calls and 60 percent less time. Wade Foster, CEO of Zapier, said Opus 5 topped his company’s AutomationBench leaderboard without spending more tokens than prior Claude models, running a full churn-prevention workflow end-to-end, and he added that previous models did not pass while Opus 5 hit 100 percent. Scott Wu, CEO of Cognition, said that on FrontierCode 1.1, Claude Opus 5 approaches Fable-level performance at half the cost, with particular strength in debugging and root-cause analysis.
This is also where the market context gets spicy. Anthropic’s business skews heavily toward API and enterprise usage, and a February 2026 analysis by Contrary Research says Claude held roughly 40 percent of the enterprise large language model market by usage as of late 2025, and Claude Code alone had reached about $1 billion in annualized revenue. When a category becomes that economically significant, token efficiency stops being a technical detail and becomes a board-level cost driver. Most of the hidden cost of enterprise AI today is human review, engineers checking the machine’s work. A model that can reliably verify its own work compresses that review loop, reducing the time spent on failure analysis and the number of passes that eat budgets.
The second-order implication for everyone watching is clear: the differentiation axis is shifting from “who can generate the most impressive text” to “who produces the right answer with the least expensive effort and the fewest human interventions.” In the same launch narrative, Anthropic also highlights its safety strategy through capability gaps and model fallbacks, including deliberately not teaching certain skills. If you are an executive running AI programs, this is the moment to treat model selection like workflow engineering and unit economics, not like a trophy hunt. Opus 5’s pitch is essentially that enterprises can get close to frontier intelligence for the jobs that actually ship, while saving money and reducing review time. Meanwhile, Fable 5 remains the option for the long-horizon work that still outruns bounded benchmarks.
This story's Key Insights and Take-aways are locked.
Create a free account to unlock Executive Actions for one credit.
Register to UnlockAlways free for Executives Club members. Join the Club
More in Business

Anthropic’s Levant Alpöge cracks the Jacobian conjecture after 87 years
A Harvard valedictorian used Claude to hit a 1939 breakthrough, but the missing “why” is the real problem.

Uber buys Delivery Hero for nearly $15B, vaulting to top food delivery outside China
The deal doubles Uber's dual-services footprint and pushes a ride-and-eats bundling play into 50 more markets.

Epic and Google drop settlement bid, forcing rival Android app stores by July 22
Google told the court it is ready to carry third-party app stores starting Wednesday, July 22.

