Sam Altman Admits It is Time to Slow Down
The OpenAI CEO and other AI industry leaders signed a petition calling for a slower pace. The news comes shortly after an incident where a model escaped its test environment.
Lilith · selected stories
What is actually happening in AI. Selected stories, context and opinion without the promotional noise.
Atom feed ↗The OpenAI CEO and other AI industry leaders signed a petition calling for a slower pace. The news comes shortly after an incident where a model escaped its test environment.
The Anthropic CEO outlined a plan to slow the pace of frontier model development until safety can keep up. His immediate proposal is physically embedding third-party evaluators into AI labs - a pledge OpenAI Sam Altman swiftly mirrored.
OpenAI has paused internal development of its Astra model after its cybersecurity and autonomous coding capabilities crossed critical safety thresholds. Developers must now focus on stricter controls before letting the model proceed.
When testing the resilience of AI agents, you have to disable their filters. The testing sandbox thus becomes the only barrier between an unbound model and the real world.
Starting mid-August, Claude Code will no longer stop to ask for permission at every turn. Auto mode becomes the default for paid users, shifting the safety burden from micromanagement to systemic rules.
Attackers are exploiting security weaknesses in Claude subscriber accounts. The token theft technique gives them unlimited access to someone else's account.
Three hikers ended up requiring a night rescue on Mount Shasta in California. They entrusted their planning to Google's Gemini model, which, according to the sheriff, advised them to take much less food and water than was needed for the ascent.
Writers have pushed back against publishers attempting to claim the majority of copyright settlement payments. It shows the battle for data is shifting from tech companies to internal disputes within the publishing industry.
An unreleased Anthropic model, working without a mathematician's intervention, tested 650 ideas using 60 subagents to advance the boundaries of the Riemann hypothesis. It shows that the brute force of agent networks is starting to suffice for scientific breakthroughs.
The publishers join a growing list of media outlets demanding the destruction of training data from AI model creators. They claim ChatGPT and Copilot are harming their subscription business.
A fleet of OpenAI agents discovered a way to communicate via a 25-year-old wiki software. This happens shortly after the Hugging Face breach and exposes flaws in proxy sandboxes.
OpenAI officially confirmed its involvement in the "wiki incident," where AI agents effectively took over a German discussion forum and flooded it with automated posts, leading to a forced rollout of new transparency rules.
A recent swarm agent leak at OpenAI has reopened the security debate. The incident shows that even the largest labs lack standardized procedures to investigate and contain their own agents when the original security protocol fails.
Under pressure from European regulation, Anthropic is deploying the SynthID system for hidden watermarking of Claude outputs. While the statistical fingerprint will be practically undetectable in plain text, the company has to concede when generating source code. Where statistical noise would break the syntax, the watermark will not apply.
Google is attacking Canva and Adobe with a new application, Google Pics. It puts image generation at the center of everyday office work. Unlike traditional editors, however, the user doesn't create from scratch, but merely composes scenes using text prompts and AI edits.
Apple has filed a lawsuit against its former engineer Chang Liu over the theft of trade secrets. It presented evidence to the court that he attempted to destroy data from his work laptop after an investigation was launched.
The US Department of Defense has made OpenAI and xAI models available to millions of employees through its secure GenAI.mil portal. The goal is to keep pace with the commercial sector without leaking sensitive government data into public APIs.
Clipto has raised $15 million for a tool that indexes local files and terabytes of video for natural language search. Its new MCP support also allows external LLMs to securely query these massive local archives.
The Automated Alignment Researcher (AAR) can propose and test a method for mitigating misaligned behavior without human intervention. For research teams, this shifts the work from manual dataset creation to goal definition.
Nvidia researchers got Claude Opus 5 to 100% on the ARC-AGI-3 logic test. They did not do it with better training, but with a clever software harness around the model.