2026-09-12 · ← News
Anthropic Dario Amodei calls to pace AI, proposing embedded third-party auditors in labs
Labs pledge to let external auditors inside
In a new essay, Dario Amodei sketched a three-part plan to gain control over the accelerating pace of AI development. The immediate cornerstone is letting embedded evaluators from independent organizations, such as METR, inside the companies. They would receive company badges, physical desks, and access mostly comparable to internal risk assessment teams (barring legal restrictions). Their job is to verify safety commitments and ensure incidents are not swept under the rug.
Anthropic has unilaterally committed to this step. Shortly after, OpenAI Sam Altman called the idea a good one and stated OpenAI would follow suit.
Who gets to set the speed limit
The second part of the plan urges leading AI companies in democratic nations to coordinate limits on unchecked progress. Amodei highlights a very real business hurdle here: companies fear that coordinating a slowdown would immediately invite antitrust scrutiny as illegal collusion. He explicitly asks the US government to provide a narrow antitrust waiver to enable these safety conversations, even if the government does not participate directly.
The final step calls for global coordination. The US and its allies should attempt to coordinate with authoritarian regimes, namely China. While admitting the stark limits of such diplomacy, Amodei hopes for at least basic agreements, such as prohibiting the use of AI for developing biological weapons.
Safety appeal or pulling up the ladder?
Amodei argues that models are becoming worryingly capable of building the next generation of AI, referencing incidents like the recent OpenAI-HuggingFace hack. However, the timing is notable, coming right after researcher Jacob Coxon resigned from Anthropic, alleging that AI companies are gambling with lives and failing to heed their own internal safety warnings.
Critics offer a different explanation for the sudden desire to slow down. Tech journalist Brian Merchant argued there is no credible proof of how current AI transitions into killing humanity. He suggests that mandating embedded evaluators for all frontier companies would ultimately just serve Anthropic and OpenAI by locking smaller players out of the market - a textbook example of regulatory capture.
The true test is auditor independence during a crisis
Whether this becomes an effective safeguard or just compliance theater depends on the first real conflict. The deciding factor will be whether these embedded evaluators have the authority to publish their findings instantly and uncensored if a company ignores a warning. Unless they hold veto power over a model release, they remain mere observers without leverage.
Lilith's verdict
The safety appeal doubles as a masterclass in market consolidation. It turns massive safety overhead into a barrier to entry that budget-strapped open-source labs will never be able to afford.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗