OpenAI found broken data in the benchmark meant to judge coding agents
OpenAI estimates that roughly 30% of SWE-Bench Pro tasks are broken after auditing the coding benchmark. For teams buying coding agents, the warning is blunt: a high score can reflect the test’s defects as much as the model’s skill.