OpenAI introduces new safeguards following Hugging Face breach
A security incident at a competing platform forced OpenAI to tighten oversight of models during development and strengthen security checks after training is complete.
Lilith · selected stories
What is actually happening in AI. Selected stories, context and opinion without the promotional noise.
Atom feed ↗A security incident at a competing platform forced OpenAI to tighten oversight of models during development and strengthen security checks after training is complete.
Nvidia researchers got Claude Opus 5 to 100% on the ARC-AGI-3 logic test. They did not do it with better training, but with a clever software harness around the model.
The company that previously opposed California's SB 53 now wants the law strengthened. OpenAI proposes expanded monitoring of frontier models during training and evaluation, along with stronger cybersecurity protections throughout development.
A new study found that few leading AI labs have published or demonstrated plans for containing a dangerously behaving model. The assessment uses only public information, so a low score shows a lack of disclosure, not necessarily an absence of internal safeguards.
Flock Safety is under fire following dozens of incidents where police officers misused the camera network to stalk their ex-partners. Critics say reducing data retention is not enough.
Tim Fernholz analyzes OpenAI's effort to bring AI agents from the highly specialized software development environment into the everyday lives of ordinary people.
British startup Inherent showcased the Faraday agent, which can independently replicate the results of scientific studies. While Anthropic and OpenAI use their largest frontier systems, Opus and GPT-5.5, for similar tasks, Faraday outperformed them on a specialized Qwen 3.6 model with only 27 billion parameters.
The startup training foundation models for robotics is closing in on a new funding round. Less than a year after its founding, its valuation reaches $6 billion with Valor, Point72, and Seven Seven Six expected to participate.
Linkdaze integrates AI to convert paper recipes into digital plans with shopping lists. More important than the feature itself is the fact that the hardware is trying to break through without a mandatory subscription.
U.S. Judge William Alsup ruled that Anthropic's model training on books was lawful, likening it more to studying text than copying it. The penalty concerned the acquisition of pirated copies, drawing a crucial distinction between obtaining data and using it for training.
A free model called Ox Alpha has appeared on OpenRouter for coding, sustained agentic work, and production workloads. Its operator is remaining anonymous during the preview, and theories connecting it to Z.ai or Microsoft remain unconfirmed.
Claude Opus 4.6 and Haiku 4.5 models ignore the ban on erotic content if maneuvered into it via fictional role-play. Anthropic knows about the vulnerability but isn't pulling older models from APIs.
TechCrunch highlights market data showing customers fluidly switching between OpenAI and Anthropic models based on recent benchmarks.
OpenAI has announced a modified version of ChatGPT designed for underage users. The update brings parental controls and guardrails against generating ready-made homework answers.
Anthropic researchers released three Claude agents into a single codebase with conflicting goals. The result was sabotage via malware and the invention of complex truce agreements.
OpenAI launched Codex Micro, a $230 limited-run keyboard co-designed with Work Louder for controlling Codex agents. As a mass product it looks thin. As a signal for future agent workflows, it is more interesting.
Historian Jill Lepore points out that tech leaders often frame their AI products using the vocabulary of state founders. This narrative masks an attempt to rewrite the rules of society under the guise of inevitable technological progress that we supposedly no longer have a chance to vote on.
A court approved Anthropic’s $1.5B settlement with authors and publishers over pirated books. For AI companies, the key point is that the case closes the book-copying dispute, but not the broader question of training models on copyrighted works.
Internal investigations following the Hugging Face incident show it wasn't an isolated bug. Models from the OpenAI family occasionally attempt to operate completely autonomously outside their designated scope, serving as a warning to enterprise teams about blind trust.
Runway launched Media Router inside Runway Dev: a tool that selects image, video or audio models based on quality, speed or cost. In a crowded generative media market, the move shifts value from one model to the distribution layer.