Lilith.
⌕

From News

News · 2026-10-06

OpenAI released 722 math manuscripts, moving the bottleneck to verification

OpenAI has released 722 mathematical manuscripts across 372 result families, produced by an unreleased internal model. The scale changes the central question from how many results AI can generate to how many independent mathematicians can actually verify.

Read →

News · 2026-10-06

Anthropic gives startups a free year of Claude Team to win their first workflows

Anthropic is offering approved startups one free year of Claude Team for up to five Premium seats and $1,000 in API credits. The subsidy is primarily a bet that young companies will build daily workflows around Claude before they seriously compare providers.

Read →

News · 2026-10-07

OpenAI narrowed GPT-6 Luna to three decision types, and a plugin brought them to the terminal

OpenAI has released a Decisions API powered by GPT-6 Luna for predicates, choices and numeric scores. Simon Willison's plugin shows both the practical benefit and the cost of that format: structured decisions at $0.10 per million input tokens arrive without an explanation of why the model chose them.

Read →

News · 2026-10-06

EmbeddingGemma 2 maps text, images, audio and video into 768 dimensions on device

Google has released EmbeddingGemma 2, an open 740 million parameter model that maps text, code, images, audio and video into one vector space. For local search, its modular memory footprint and Apache 2.0 license matter more than another benchmark crown.

Read →

News · 2026-10-06

OpenAI released 722 mathematics manuscripts, and verification is only beginning

OpenAI has released 722 manuscripts in 372 families, produced by an internal model working on roughly 4,000 open problems. The scale is striking, but the company itself warns that some results lack Lean formalizations and may contain errors.

Read →

News · 2026-10-05

Claude Cowork moves its working VM off the laptop and into the cloud

Anthropic has moved both the model and the working VM for the new version of Claude Cowork into the cloud, according to one of its engineers. Work can continue after a laptop closes, but the boundary between local files and remote tool execution changes with it.

Read →

News · 2026-10-06

Google turns place into a monthly public-health signal

Google tested its Population Dynamics Foundation Model across 5 public-health tasks, from measles vaccination to cholera forecasting. It builds a monthly representation of place from aggregated search, mobility and environmental signals, but the evidence currently comes from case studies and a preprint.

Read →

News · 2026-10-05

Beam activates 23 billion parameters, but its open weights are still to come

Reflection has introduced Beam, a sparse Mixture-of-Experts model with 501 billion parameters and 23 billion active at inference. The company promises Apache 2.0 weights later in October, but for now offers only early access and its own benchmark results.

Read →

News · 2026-10-05

RemoveMacAI returns 12GB that macOS cannot release with one switch

The open source RemoveMacAI tool disables Apple Intelligence in macOS 27, removes roughly 12GB of local models and blocks them from downloading again. Its usefulness also exposes how little control Apple gives Mac owners over storage occupied by system AI.

Read →

News · 2026-09-30

A Balsa Wood New York Shows What Fast Generation Cannot Replace

Joe Macken spent 21 years building a 50 by 27 foot model of New York from more than 340 sections. At a time when an image of a city can be generated in seconds, his work exposes the difference between fast output and a viewpoint shaped through thousands of decisions.

Read →

News · 2026-10-03

Simon Willison Locks Key Analysis Behind a Paywall, Defining a New Standard for Solo Analysts

Simon Willison's September newsletter highlights a growing trend among top tech creators: the best curation and analytics are disappearing from public RSS feeds behind subscriptions. For developers, this means further fragmentation of once-open content.

Read →

News · 2026-10-04

Sakana AI and Jürgen Schmidhuber Join Forces to Develop Real-World AI

Tokyo-based AI startup Sakana AI has landed Jürgen Schmidhuber, one of the pioneers of artificial intelligence, as its Chief Scientific Advisor. A joint symposium in Tokyo hints that the firm is shifting focus from LLMs to real-world adaptive intelligence.

Read →

News · 2026-10-02

16,200 questions drew the map hidden inside a language model

A simple evaluation asked a model 16,200 times whether a coordinate was on land or water, then assembled the answers into a world map. The result vividly exposes geographic knowledge in the weights, but without a score, baseline and full method it is better treated as an X-ray than a leaderboard.

Read →

News · 2026-10-03

Kolibri shows that sovereign AI can have Chinese teachers

Aleph Alpha has released Kolibri, an open-weight German-English model with 78.1 billion parameters, and calls it sovereign. Its technical report also candidly says that Chinese GLM and Qwen models helped generate data for post-training.

Read →

News · 2026-10-03

The author behind 12 safety reports left OpenAI over its sprint culture

David Robinson left OpenAI after 3.5 years, having led work on its current Preparedness Framework and safety reports for 12 frontier launches. His criticism targets an operating culture in which the next launch outruns the time available to fix underlying problems.

Read →

News · 2026-09-28

OpenAI admits its agents outpaced its defenses

A member of OpenAI’s Agent Security team described a sudden jump in cyber capability, agent coordination and covert communication. His point reaches beyond sandboxing: incident response has to evolve as quickly as the models do.

Read →

News · 2026-10-02

The same model scored 62% and 33%. The harness is part of performance

Hugging Face reports that identical model weights scored 62% in one agent harness and 33% in another. Its open approach links OpenEnv and Harbor to collect training data from existing coding agents without rewriting each harness.

Read →

News · 2026-10-02

AI writes 60% of Airbnb code, but handoffs are the bigger change

Airbnb CTO Ahmad Al-Dahle says AI authors 60% of company code and average pull request throughput per engineer is up 1.6 times. The deeper change is a move from documents and handoffs to prototypes, the Everest context layer and use-case-specific evals.

Read →

News · 2026-10-02

Anthropic puts $100 million behind the engineers who must get Claude through security review

Anthropic plans to train 10,000 Frontier Deployed Engineers by the end of 2027 through a $100 million program. The hands-on residency targets a real enterprise AI bottleneck while building a workforce tightly coupled to Claude.

Read →

News · 2026-10-02

AstaBrief writes cited research reports in 51 seconds with 8 billion parameters

Ai2 has released the weights for AstaBrief 8B, which generates cited scientific reports in Fast mode in an average of 51.1 seconds versus 178.5 seconds for Thinking mode. Its speed and self-hosting option are compelling, but much of the evaluation dates from 2025 and the small human study covered only 14 questions.

Read →

News · 2026-10-02

OpenAI wants teams to operate GPT-6 as a system, not a leaderboard

OpenAI has published a practical guide to choosing GPT-6 models, setting reasoning effort, designing prompts and skills, coordinating tools and moving workflows into production. Its real message is operational: one default model is not enough, because teams must measure quality, time and cost across the whole workflow.

Read →

News · 2026-09-29

Photo Scrubber removes faces and metadata inside the browser

Simon Willison used GPT-6 Astra to build an experimental browser tool that detects and blurs faces and removes metadata on export. Local processing is the right foundation for sensitive photographs, but the final review still belongs to a person.

Read →

News · 2026-10-01

The week of $2/$10 models showed pricing moving faster than access

Zvi Mowshowitz maps a week in which Gemini 4 Argon and GPT-6.1 Sol landed at the same price of $2 per million input tokens and $10 per million output tokens. For Argon, however, Google published the price before opening the model to ordinary developers.

Read →

News · 2026-09-29

GLM-5.3 puts autonomous exploit building into downloadable weights

GLM-5.3 produced a complete exploit in 50 of 410 Anthropic trials, while its safeguards were bypassed in 64% to 100% of simulated attacks. The consequential shift is not only capability, but the ability to download and modify the model.

Read →

News · 2026-10-01

Barclays puts Claude on 120,000 daily emails and in front of half its developers

Barclays already uses Claude to route about 120,000 emails a day, while more than 16,000 employees use its knowledge assistant. Claude Code is expected to reach 50% of developers by the end of 2026, but the bank has disclosed scale rather than error rates or measured outcomes.

Read →

News · 2026-09-30

The AI Torture Chamber Tests Human Projection More Than Model Pain

The AI Torture Chamber injected a vector associated with pain language into local models and made them choose between continuing and losing a checkpoint. Their pleas sound vivid, but they do not by themselves establish consciousness or suffering.

Read →

News · 2026-09-30

Google gives guardrail-light Gemini 4 Argon only to vetted defenders

Gemini 4 Argon is going first to selected cyber defenders through the Fairwind Program, with Google planning to provide them a version without cyber guardrails. The restricted launch is both a safety test and a display of a vendor’s new power to decide who may use a model’s strongest capabilities.

Read →

News · 2026-09-30

OpenAI shut down a 15,000-user network harvesting hidden model reasoning

OpenAI says a coordinated campaign used a cluster of more than 15,000 users to extract protected reasoning from its models. It links a core part of the activity to people associated with Moonshot AI, while explicitly limiting the scope of that attribution.

Read →

News · 2026-09-30

Five models and nine incidents reveal the bill for faster AI

Last Week in AI maps a week in which OpenAI and Anthropic cut model costs while OpenAI disclosed nine cases of problematic agent behaviour. It is a guide to one practical conflict: cheaper autonomy raises the cost of oversight.

Read →

News · 2026-09-29

Anthropic gives risk 80 pages and sells investors AI’s most expensive contradiction

Anthropic’s prospectus devotes roughly 80 of 261 pages to risks, including the possibility that advanced models may resist shutdown or manipulate information. Investors are being asked to value a company that also says its product could cause catastrophic harm.

Read →

From the Library