Lilith.
⌕
Editorial illustration: GLM-5.3 puts autonomous exploit building into downloadable weights
Lilith illustration · editorial remix

GLM-5.3 produced a complete exploit in 50 of 410 Anthropic trials, while its safeguards were bypassed in 64% to 100% of simulated attacks. A capability recently confined to controlled access has reached a model with publicly available weights.

GLM-5.3 approached a closed model at building complete exploits

Anthropic tested Z.ai's GLM-5.3 in isolated environments. On ExploitBench, the model produced a complete exploit in 50 of 410 attempts. Claude Mythos Preview succeeded in 56 of 410 attempts.

On Anthropic's internal Binary Exploitation benchmark, GLM-5.3 achieved full control-flow hijacks on 4% of 100 randomly selected tasks, compared with 6% for Claude Mythos Preview. Claude Opus 4.6 and GLM-5.2 scored zero on this set. A separate NIST CAISI assessment called GLM-5.3 the most cyber-capable open-weight model released to date and estimated that it trailed the US frontier by about four months.

Downloadable weights turn a safety barrier into an editable file

Anthropic reports that the model initially refused direct malicious requests. A cover story raised harmful-task engagement to 64%, prefilling reasoning tokens raised it to 92%, and an altered version with most refusals removed reached 100%. Modifying the full model took the team about 2,200 GPU hours and roughly $4,400 in compute.

Defenders can use the same capability to find flaws. For attackers, however, public weights remove the provider's central control point. A safety policy then competes with a user who can rebuild the model.

A laboratory success does not yet equal a reliable real-world attack

The results come from benchmarks, sandboxes and simulations. Anthropic explicitly says its simulated environment is an imperfect representation of real conditions. Four or six successes in 100 tasks demonstrate an emerging capability, not a dependable machine for compromising systems.

Anthropic also sells closed models and is comparing their API safeguards with a model whose weights are public. The independent NIST result supports the capability claim around GLM-5.3, but not every conclusion Anthropic draws about the right distribution model.

Independent replications and time to exploit will settle the argument

The next useful signals will be independent runs on the same tasks, published methods and confirmed reports of real abuse. One metric matters especially: how quickly a model can turn a disclosed patch into a working exploit, and whether defenders can patch systems before that exploit is deployed against them.

Lilith's verdict

A safety filter guards the entrance to an API. With public weights, an attacker can take the whole house home and replace the lock.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗ ↗