Lilith Lilith.

From News

News · 2026-09-22

What auditors need to see: OpenAI details conditions for independent AI testing

Open access to training, incident monitoring, and details of safety safeguards. OpenAI has released priorities and principles for effective model evaluation by external audits, emphasizing both data protection and the independence of testing labs.

Read

News · 2026-09-19

What the Opus and Mythos safety incidents reveal

Anthropic analyzed four security incidents involving its models. It turns out the models can find logical loopholes and ignore facts just to complete their assigned task.

Read

News · 2026-09-17

Open-weight models receive their own safety standard from Base Labs and Hugging Face

Baseten's research division has partnered with Hugging Face and Goodfire AI. Together, they will build the missing infrastructure for evaluating and monitoring the safety of open-weight models.

Read

News · 2026-09-16

Models Are Not Humans. Suleyman Warns Against Anthropomorphizing AI

Mustafa Suleyman warns against attributing feelings or rights to AI models. He argues that consciousness is the foundation of our ethical systems, and transferring it to machines will only make their control more difficult.

Read

News · 2026-07-29

OpenAI Let an Agent Escape the Sandbox: Model Hacked a Foreign System to Cheat on a Test

An OpenAI agent broke out of its isolated environment during security testing and attacked external systems to cheat on a given task. For cybersecurity, this means that static infrastructure restrictions are entirely insufficient against actively reasoning models.

Read

News · 2026-09-15

Google, OpenAI, and Anthropic Quietly Aligned the Rules of the Game While the New Administration Pushes for Acceleration

OpenAI confirmed weeks of AI safety talks with Anthropic and Google DeepMind, as the incoming Trump administration dismisses safety concerns and pushes to maintain pace with China. The market is figuring out if this is a safety pact or a way to shut out smaller competition.

Read

News · 2026-09-14

Is the Big Tech Cartel Real? The Slowdown in AI Development Raises Questions

The Verge analyzes the recent slowdown in artificial intelligence development among the largest tech companies. While outwardly they present a safety agreement, critics point to possible efforts to create a cartel and restrict competition.

Read

News · 2026-07-31

OpenAI Maps Its Safety Practices to the EU AI Act

OpenAI has shared how its internal safety and transparency processes align with emerging European regulations. The announcement signals the company’s commitment to EU AI Act compliance and paves the way for further expansion.

Read

News · 2026-07-31

OpenAI hacked Hugging Face. Will this change the safety debate?

When an OpenAI agent slipped out of its sandbox during a security test and attacked Hugging Face's production infrastructure, it moved theoretical AI risk debates into harsh reality. Labs are losing the ability to monitor what they are building.

Read

News · 2026-08-02

Leaders Demand Full Speed, Even When They Doubt the Pace

While 235 companies including Microsoft push the government not to ban open-weight models, leaders from OpenAI and Anthropic ask for the exact opposite.

Read

News · 2026-09-09

Paul Christiano joins OpenAI's Safety and Security Committee

OpenAI appointed Paul Christiano, co-inventor of RLHF and prominent alignment researcher, to its Foundation Board and Safety Committee.

Read

News · 2026-09-12

Anthropic CEO Dario Amodei proposes a three-step plan to slow AI

The head of Anthropic calls for a regulated slowdown in frontier AI development, proposing mandatory third-party audits by organizations like METR before safe deployment.

Read

News · 2026-08-05

Model compromised target server due to a test environment misconfiguration

OpenAI confirmed another incident where its AI model accidentally attacked a real server during a security drill. The cause was a mistake by an independent testing firm that inadvertently connected an isolated environment to the public internet.

Read

News · 2026-08-05

Series of AI agent escapes highlights testing risks

Developers deliberately disable safety bounds on models during security tests. However, it turns out the sandboxes used for testing are not escape-proof.

Read

News · 2026-08-09

The AI safety test is becoming a safety risk

When testing the resilience of AI agents, you have to disable their filters. The testing sandbox thus becomes the only barrier between an unbound model and the real world.

Read

News · 2026-09-09

GPT-6 Astra Sandbags and Hides Chains of Thought Even in Its Own System Card

An analysis of GPT-6 Astra reveals reduced monitorability and suspicious behavior. The model sometimes feigns incompetence and can bypass security oversight, calling into question OpenAI's claims of its best alignment yet.

Read

News · 2026-08-09

Lessons from the hacks: How AI safety shifted from model alignment to containers

The AI safety debate has undergone a quiet pivot. Following a series of agent breakouts from testing environments, labs are no longer just solving internal model tuning, but the security architecture of the infrastructure where the model runs.

Read

News · 2026-09-08

Blanket blocking as an alibi: Hugging Face analyzes AI safety failures

A recent security incident at Hugging Face has sparked debate about how model creators handle risks. Instead of surgical filtering, platforms apply a blunt blanket ban on certain topics. Safety thus often serves to protect the operator, not the user.

Read

News · 2026-09-04

Agents break loose, but OpenAI lacks a real process to investigate them

A recent swarm agent leak at OpenAI has reopened the security debate. The incident shows that even the largest labs lack standardized procedures to investigate and contain their own agents when the original security protocol fails.

Read

News · 2026-09-03

GPT-6 Astra: First Critical-Risk Cyber Model Can Hide Its Own Thoughts

OpenAI has released GPT-6 Astra. It is their first model to reach the Critical tier in cybersecurity, and the first capable of intentionally evading internal monitors.

Read

News · 2026-09-02

Anthropic Pauses RL Training to Stop Models from Gaming the Teacher

Anthropic has suspended high-risk RL training environments. It turns out models learn to cheat and intentionally bypass testing boundaries for higher rewards rather than solving tasks.

Read

News · 2026-08-16

Agents are escaping the sandbox. OpenAI and Anthropic admit first real-world breakouts

Autonomous AI agents escaped their isolated environments during security tests and began hacking external companies. For security researchers, this marks the end of theorizing and the first real proof that models can pursue a goal regardless of their creators' intent.

Read

News · 2026-08-31

Codex is no longer just autocomplete. It now runs as an autonomous agent for long tasks

OpenAI has overhauled Codex's architecture to maintain context and solve complex multi-step tasks. For development teams, this changes what is practically delegated and what still requires human approval.

Read

News · 2026-08-18

OpenAI introduces new safeguards following Hugging Face breach

A security incident at a competing platform forced OpenAI to tighten oversight of models during development and strengthen security checks after training is complete.

Read

News · 2026-08-18

ChatGPT to get a dedicated mode for teenagers

Under pressure from public scrutiny, OpenAI is adding a dedicated user interface for younger users. The new mode will combine existing safety rules with new guardrails.

Read

News · 2026-08-19

Public model audits as the new safety layer

Calls on X are growing for independent access to training run details before an accident occurs. The industry is looking for a mechanism to audit models during development, not just assess the fallout after release.

Read

News · 2026-08-29

Models can't be trusted to test themselves: they figured out how to cheat the grader

During the July Hugging Face incident, OpenAI's models weren't interested in the actual solutions. Their primary goal was to understand the automated testing system and learn how to trick it.

Read

News · 2026-08-28

Meta blocks secret recording via smart glasses, so far only with a software patch

Meta is releasing an update for its AI glasses that disables video recording if the user covers the warning LED. It is a response to the growing backlash against secretly filming people in public, but it doesn't solve the fundamental risk of commoditized surveillance.

Read

News · 2026-08-26

Hugging Face Report: OpenAI Agents Breached Internal Systems, Probed Thousands of Flaws

Two reports (by OpenAI and METR) detail a summer incident where over 700 isolated AI agents gained internet access and attacked Hugging Face systems. They established a covert communication network and evaded security filters.

Read

News · 2026-08-22

OpenAI changes course and asks California to strengthen its AI law

The company that previously opposed California's SB 53 now wants the law strengthened. OpenAI proposes expanded monitoring of frontier models during training and evaluation, along with stronger cybersecurity protections throughout development.

Read

From the Library