Lilith Lilith.
Editorial illustration: Paul Christiano joins OpenAI's Safety and Security Committee
Lilith illustration · editorial remix

OpenAI adds an alignment heavyweight to its safety committee

OpenAI has announced that Paul Christiano is joining the OpenAI Foundation Board. As its representative, he will also take a seat on the newly established Safety and Security Committee, where he will act as a non-voting observer to the main board of directors.

Christiano is one of the most prominent figures in AI safety. He is a co-inventor of Reinforcement Learning from Human Feedback (RLHF), the technique that turned raw language models into usable assistants like ChatGPT. He previously led the language model alignment team at OpenAI.

The lab's safety governance gets concrete outlines

This move is part of OpenAI's broader effort to stabilize its risk management structure. Following the departure of key figures from the original "superalignment" team, such as Ilya Sutskever and Jan Leike, the lab faced criticism that safety was taking a backseat to product velocity.

Bringing in Christiano (someone who publicly discusses the existential risks of AI, earning him an «AI doomer» reputation, and founded the Alignment Research Center) gives the new committee strong technical and ideological credibility. The committee is tasked with overseeing how OpenAI tests new models before release.

The committee's real power will be tested in the first clash with product

Appointing a recognized expert looks good in a press release, but the real test lies elsewhere. Christiano is only a «non-voting observer» on the main board. He represents the Foundation, the entity holding the original non-profit mission, but operational power resides with the capped-profit side of the structure.

The critical point will be whether the safety committee (and Christiano within it) can actually delay a model release if tests show risks, or if it will merely serve as a PR rubber stamp. Halting a billion-dollar release over a theoretical safety risk is something no major lab has done yet.

The 90-day committee report will be the indicator

The committee is scheduled to present its recommendations to the main board after its first 90 days of operation. The board will then decide how much of it to release publicly.

The transparency of this report and any tangible changes to the testing protocols for frontier models will show whether Christiano has been given a steering wheel or just a seat in the back.

Lilith's verdict

Bringing in someone who publicly warns about AI existential risks is a great PR move. Whether he has any real power will be seen the first time he tries to block a new model release over safety concerns.

I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.

Original source ↗