2026-09-15 · ← News
Attack report shows the limits of bad intentions: Claude blocks attackers
Surviving in a hostile environment
A new analysis by Zvi Mowshowitz examines a recent report from Anthropic detailing attempts to misuse the Claude model. Apparently, attackers are not having an easy time. Many actors try to convince the LLM to assist them with fraud, cyberattacks, or other illegal activities. However, Anthropic actively disrupts these attempts, and according to available data, most of them fail.
Barriers are holding for now
This situation is mildly optimistic for security teams. It shows that despite theoretical concerns about the misuse of advanced models, current mitigations at the API and model level provide a solid defense against common attacks. If Anthropic's report reflects the true state of affairs (and not just carefully selected examples from the tip of the iceberg), it means that turning an LLM into an obedient accomplice for organized crime is harder than it seemed.
The invisible threat below the surface
There remains, however, a massive blind spot. The report only shows the attacks that Anthropic detected and stopped. We know nothing about those who were smart enough to evade detection, or those using open-source models without any vendor oversight. The true test of security is not the blocked attempts of amateurs, but the sophisticated attacks of state actors, which we usually do not read about in such reports.
An arms race at the prompt level
The key will be to watch how attacker techniques evolve. When standard filter bypasses (jailbreaking) become too difficult, attention will likely shift to attacks exploiting vulnerabilities in the architecture of the agents entrusted to models like Claude, and the manipulation of external tools.
Lilith's verdict
Reports about intercepted attacks are PR material. The real threat isn't an amateur asking how to make explosives, but a professional who downloaded uncensored weights long ago and isn't asking anyone.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗