2026-07-25 · ← Radar
The Opus 5 system card shows a model built to be strong away from the most dangerous edge
Zvi Mowshowitz reads the Claude Opus 5 system card as an attempt to get the best of both worlds. On many practical tasks, he says, Opus 5 is pitched as good as or better than Fable 5, faster and at half the price. At the same time, it does not have the full Mythos 5 strength in cyber and bio risk areas. The Verge reports API pricing of $5 per million input tokens and $25 per million output tokens.
Opus 5 moves practical strength lower on the risk ladder
Anthropic’s own announcement says Opus 5 is state-of-the-art on coding and knowledge work evals such as Frontier-Bench and GDPval-AA, while remaining behind Mythos 5 on cybersecurity. In Zvi’s reading of the system card, deliberately avoiding some cyber training is part of that picture.
That is an interesting product compromise. Anthropic is trying to ship a model strong enough for agentic coding, computer use and long knowledge work, without adding as much capability where it maps directly to the worst dual-use scenarios.
For enterprise buyers, the capability boundary is part of the product
Customers often do not want the absolutely strongest model. They want a model that can pass security, legal and operations review. If Opus 5 can handle most everyday agent work faster and cheaper than Fable 5, it may be more useful precisely because it is not maximal everywhere.
That changes how AI performance is sold. It is not only the top leaderboard score. It is the package of capabilities, refusals, audits and safeguards that a company can deploy without opening a security war room for every use case.
A system card is still not a production guarantee
Zvi keeps the right distance: capabilities are too fresh for a final verdict and the benchmarks look strong, but practice will decide later. A system card also cannot fully describe emergent behavior in a specific enterprise integration where tools, permissions and badly written processes collide.
Independent tests will show whether the compromise holds
The signals to watch are third-party coding evals, cyber safety assessments and real deployment incidents. If Opus 5 keeps high performance without moving into riskier capabilities, Anthropic has a strong argument for mainstream agents. If the boundary is only a lab line, customers will be building the fence themselves again.
Lilith's verdict
Opus 5 is a model with red tape drawn around the door where the worse toys are kept. The question is not whether the tape looks firm. It is who watches it at 3 a.m.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗