2026-09-27 · ← News
Opus 5.5 brings back what benchmarks miss: personality
Ethan Mollick describes Opus 5.5 as a return to the old feel of Claude. In his view, Opus models from roughly 4.7 through 5 lost part of their distinctive personality and resembled a diluted Fable. This is one user's experience, not a measurement. It still points at a blind spot in today's leaderboards.
Anthropic says the leaderboard tells only part of the story
Anthropic released Opus 5.5 on September 22, 2026. The company says it communicates more naturally than previous models, puts important information first and addresses criticism of Opus 5's dense writing. It also says benchmark margins have become a less reliable guide to differences in real work at current capability levels.
The release includes changes that are easier to measure. Anthropic claims 40% lower costs on typical workloads than Opus 5 and output generation that is more than 30% faster. API pricing is $4 per million input tokens and $20 per million output tokens. The model is available through Anthropic, AWS, Google Cloud and Microsoft Azure.
Model style determines the cost of human correction
Personality is consequential during long tasks. It affects whether a model admits uncertainty, how it orders an argument and how much of its output a person must rewrite. Two models with the same score can therefore impose very different costs in editor, developer or analyst hours.
Mollick's post adds a user layer to the vendor's numbers. Anthropic measures tokens, speed and task success. A user also measures whether the model remains workable through a three-hour session without constant correction of its tone and judgment.
The sense of a comeback still rests on anecdotes and vendor claims
Claude-like has no standardized metric. Mollick did not publish a blind test or a prompt set, while Anthropic quotes its early testers. The improved style could reflect a genuine model change, a system prompt effect or simply the novelty of a fresh release.
More natural communication can also diverge from accuracy. Fluent answers are easier to read, but fluency alone does not show that a model corrects errors, follows instructions or resists sycophancy more reliably.
Blind preferences and long sessions will show whether Claude is truly back
Repeated blind comparisons on identical tasks and real working sessions are the useful signals now. Teams should track human edits, tone stability across dozens of turns and whether the model changes its position when presented with evidence.
If the preference holds across users and domains, labs will need to measure consistency of collaboration alongside capability. A benchmark will remain a scoreboard, not a complete product description.
Lilith's verdict
A leaderboard can rank models. We will know Claude is truly back when a colleague erases fewer corrections from the whiteboard after a three-hour session.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗