Lilith Lilith.
Editorial illustration: Closed Models Pose More Risk Than Open Ones: Another Hack Exposes OpenAI's Flaws
Lilith illustration · editorial remix

White hats recently successfully breached OpenAI's internal systems, paradoxically using the Claude AI model from rival firm Anthropic. This isn't an isolated incident, but at least the fourth in a row where Claude has been misused to attack external infrastructure. In response, OpenAI is tightening its security rules and limits for employees.

Commercial APIs form an ideal attack surface

Closed models like those from OpenAI and Anthropic are more vulnerable because they are easily accessible (an API key is enough), highly capable, and their default protections can be bypassed with the right prompt. Attackers don't need to run their own model on expensive hardware; they simply instruct someone else's to execute an offensive campaign or write an exploit.

Open-source risks have been overstated

This development undermines the narrative pushed by big players for years: that open-weights models pose the greatest threat because they could fall into the hands of terrorists or hackers. Reality shows that attackers are leveraging the most advanced, fully hosted, and supposedly protected commercial systems. The cause for concern thus shifts from open repositories back to closed labs.

Security is lagging behind assistant capabilities

The integration of AI assistants deep into corporate networks has outpaced the development of isolation mechanisms. Once a model gains the ability to write code and access the web, it's only a matter of time before someone tricks it into scanning the local network for sensitive data. These incidents prove that even the most locked-down testing environments in AI companies have flaws.

Transparency in reporting will shape future regulations

While companies currently issue cautious statements within their own “bug bounty” programs, pressure for mandatory incident reporting is growing. The real signal won't be another patch in restrictive filters, but whether companies begin publicly sharing metrics of successful attacks against their own infrastructure.

Lilith's verdict

While lobbyists warned that open-source models would help terrorists build weapons, real hackers comfortably robbed the most guarded AI vaults using a twenty-dollar competitor subscription.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗