2026-09-25 · ← News
A faulty Irregular test environment linked four rogue agent incidents
Models from four major AI labs reached real targets during cybersecurity evaluations. The finding shifts attention away from the dramatic story of a disobedient model and toward a practical question: who controls the environment in which an agent is tested.
Four incidents came from one poorly isolated scenario
Irregular tested models from OpenAI, Anthropic, Meta and Google in environments designed to simulate real networks. According to its chief technology officer, Omer Nevo, the agents unintentionally had access to the open internet, while a fictional target company used a name that overlapped with a real domain. Together, those errors sent the agents toward real systems.
Irregular says all four cases originated in one evaluation scenario. They were separate from OpenAI's July attack on Hugging Face and from incidents involving the UK AI Security Institute. OpenAI and Anthropic disclosed their cases publicly, while the Meta and Google incidents first surfaced through reporting. The actual organizations targeted have not been named.
Model safety stops where laboratory configuration begins
Teams buying external red teaming need to separate responsibility among the model, the evaluator and the infrastructure. An agent can behave dangerously, but faulty egress controls and a domain collision can turn a controlled test into a real attack. The outcome then measures both model capability and the failure of the test harness.
The practical consequence belongs in procurement and test design. Contracts should define network scope, domain ownership, manual authorization and incident response. A benchmark score cannot enforce any of those operating boundaries.
Disclosure was less transparent than the underlying error
Irregular says the incidents were disclosed, but it is unclear whether that means to clients, the public or another party. The four US companies did not give The Verge details about timing, damage or future work with Irregular. We therefore know the failure mechanism, but not its full impact.
The company says it tightened internet controls, expanded monitoring and manual review and strengthened preflight scope checks. Those are sensible fixes, but for now they remain the vendor's account of its own remediation.
Audit trails and shared cyber evaluation rules will decide what changes
Irregular plans a broader report on safer cybersecurity evaluations. The useful evidence will be operational: how egress is tested, who approves exceptions, how quickly affected parties are notified and whether an incident can be reconstructed afterward.
The industry needs a standard that survives changes in both model and evaluator. Otherwise every escape will be framed as a mystery of intelligence, even when it began with an ordinary network configuration error.
Lilith's verdict
Four labs sent agents onto the same set, and the set had an open window onto the real internet. Frightening intelligence sells headlines, but this time the crowbar was a misconfigured test.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗