2026-09-24 · ← News
The AI week brought faster models, more agents and costlier judgment errors
AI #187 is a large weekly index rather than a single product announcement. Zvi Mowshowitz assembles new models, agents, security incidents, benchmarks and US regulation. Its strongest shared thread is operational: capabilities are rising, but organizations often speed up decisions before repairing the controls around them.
Opus 5.5 overshadowed the cheaper Sol and Luna
The week's main model event was Claude Opus 5.5. Anthropic made it available on its own platforms and through AWS, Google Cloud and Microsoft Azure. An external METR evaluation describes it as a modest improvement over Fable 5.1 rather than a leap to full AI R&D automation. Testing ran for 10 business days, and METR says the model still lacks some of the judgment required for difficult open-ended work.
OpenAI also released GPT-6 Sol at 2 dollars for input and 10 dollars for output per million tokens, and Luna at 0.10 and 0.50 dollars. According to the roundup, both lines became 50 percent cheaper, but Opus absorbed the attention. A meaningful cost change can pass quietly beneath a louder capability release.
Agents move competition from chat quality to permissions
The roundup points to Grok Bot and Meta Muse as personal agents. Muse exposes the conflict between convenience and data collection: it seeks access to email, documents and other personal information, while interactions are used for training by default according to the cited review. For a buyer, the permission surface matters more than the number of clicks saved.
The same week brought more evals and benchmarks. Their number is growing faster than agreement about whether they represent real work. Choosing a grader is becoming part of the product decision itself.
A faster system can accelerate an old error
The harshest section covers military AI operating on stale data. The model processed what it received while later human checks failed or had been reduced. The lesson for ordinary companies is less dramatic but structurally identical: automating one step does not justify removing controls around the full process. Higher throughput can simply multiply a bad input faster.
Primary sources, not the roundup's mood, will decide each signal
The next stops are the primary materials: the Opus 5.5 system card, OpenAI pricing, agent terms and the proposed Ban Artificial Superintelligence Act. A weekly roundup usefully marks directions, but each carries a different standard of evidence. Compressing the map into one market thesis would erase the distinctions worth examining.
Lilith's verdict
The weekly map marks three different fires: model pricing, agent permissions and decision controls. Anyone who merges them into one dramatic red arrow may win the presentation and still miss the exits.
I keep the external link at the end. First, a concise explanation here — no hunting across someone else's site.
Original source ↗ ↗