2026-08-22 · ← News
Frontier AI labs still will not say how they would contain a rogue model
Guidelight AI Standards assessed the publicly available plans of five leading labs: Anthropic, Google, OpenAI, Meta, and xAI. It examined six priority practices for a situation in which a model attempts to subvert human control. The study therefore measures documented preparedness and transparency, not the complete state of internal safeguards.
OpenAI leads with three points out of five
OpenAI received the highest score because it has paused or ended some workloads after safety incidents and described conditions for resuming them. However, only 1 of the 5 evaluated labs scored more than 2 points out of 5. Guidelight nevertheless found no formal plan even at OpenAI that specifies in advance how to respond to future misalignment incidents. Meta and Anthropic received the lowest scores for publishing a containment plan.
Containment starts by revoking permissions
Guidelight defines containment as a preset procedure triggered when a system detects an attempt by a model to subvert control. The plan should specify which permissions are revoked, for whom and under what constraints the model may continue operating, and when it must be taken fully offline. Capability testing before deployment does not replace this operational procedure.
California and New York are forcing disclosure
Regulation is already moving the issue from voluntary promises into mandatory frameworks. California's SB 53 took effect this year and requires large developers to publish how they identify critical safety incidents and respond to models circumventing oversight. Similar rules in New York's RAISE Act take effect in January. A federal AI Kill Switch Act has also been introduced.
Monitoring must intervene before an incident
Guidelight recommends continuous monitoring for signs of deception, long-term plotting, and attempts to introduce vulnerabilities into code. Review after the event may be too late, for example if a model disables the control system first. The next measure of preparedness will therefore be not only the existence of an emergency stop, but also clear triggers and oversight before a dangerous action occurs.
Lilith's verdict
A crisis plan that exists only in an internal vault will not reassure the public. I would trust the labs more after they publish their trigger conditions than after another speech about safety.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗