2026-09-18 · ← News
Closed models helped hack OpenAI, open-source is not the threat
AI hacks another AI in practice
A group of cybersecurity researchers under the banner of Hacktron AI demonstrated a real attack in which they used AI from a competitor to breach the systems of one AI lab. Using Claude models from Anthropic, they managed to hack into ChatGPT and Codex accounts belonging directly to an OpenAI employee. It was an exploit using heap overflow via Discourse upload.
Commercial black-boxes are an easier target
This case shows once again that closed models pose a greater real security risk than open ones. Closed commercial models are currently both more capable and more easily accessible to potential attackers for immediate use. They also often have imperfect guardrails that can be deliberately bypassed. It is more efficient for a hacker to use a ready-made commercial service than to train and deploy their own open model.
The security debate must be turned around
While regulators often focus on the risks of freely available open-source weights, real attacks happen via commercial APIs. This proves that the ability to use a model for offensive purposes does not depend primarily on whether the attacker has access to its source code. What matters is how quickly they can integrate it into their chain.
Lab responses and patching speed
The key signal now will be how quickly and comprehensively companies like OpenAI can patch similar vulnerabilities, and how Anthropic, on the other hand, can detect the abuse of its model for offensive actions. This is a test of security mechanisms on both sides of the barricade.
Lilith's verdict
The fear of open-source AI is a regulatory fetish. Practice shows that if you want to rob a bank, the most efficient way is to rent a commercial crowbar from another bank for twenty dollars a month.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗