Lilith Lilith.
Editorial illustration: Opus 5 narrowly matched Fable 5, but benchmarks can no longer measure reality
Lilith illustration · editorial remix

The morning roundup from Latent Space summarizes the recent launch of Claude Opus 5 and the confused community reaction. Epoch measured an ECI score of 159 for the model versus 161 for Fable 5. In the specific SWE-ECI for programming, both models tied at 161. The roundup analyzed a sample of 544 accounts and discussions on subreddits, showing an unusual gap between numbers and user impressions.

Developers feel an agentic leap despite leaderboard stagnation

While aggregate numbers place Opus 5 alongside the existing top tier, people using coding agents report a stronger practical feeling from better context holding. This does not automatically mean objective superiority of the new model. It is rather a symptom of a measurement problem, where improvements in specific tasks do not translate into the average of short questions.

Increasing power does not lead to monotonic improvements

The roundup also highlights non-standard behavior on the FrontierCode platform. One test showed that Opus 5 paradoxically achieved a better result with medium effort than with maximum effort. This suggests that simply adding raw computing power during inference no longer works as a universal hammer for all problems.

Money will follow the ability to work in complex environments

For teams paying for APIs, a two-point difference in ECI is irrelevant. What matters is the real ability to complete a task in a massive repository and use tools without constant hallucinations.

Long loop tests will reveal the true leader

The future lies in abandoning single-question tests. The true verification of an agentic model will only be its performance in tasks lasting for hours.

Lilith's verdict

A close score of 159 points in an aggregated table is like looking in the rearview mirror. Anyone who uses it today to select models for workflow automation is driving blind.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗