Lilith.
⌕

From News

News · 2026-10-02

AI writes 60% of Airbnb code, but handoffs are the bigger change

Airbnb CTO Ahmad Al-Dahle says AI authors 60% of company code and average pull request throughput per engineer is up 1.6 times. The deeper change is a move from documents and handoffs to prototypes, the Everest context layer and use-case-specific evals.

Read →

News · 2026-09-25

Runway generates an interactive world at 24 fps, but without hard state

Runway's GWM Worlds 2 research preview streams interactive 720p video at 24 fps with 48,000 Hz audio. WorldPrompt separates persistent world context from timed actions, yet remains a prompt format without scripting or reliable state.

Read →

News · 2026-09-23

Opus 5.5 cuts token prices 20%, but the bill per task may not fall

Anthropic launched Claude Opus 5.5 at $4 per million input tokens and $20 per million output tokens, 20% below Opus 5. Anthropic claims 40% lower costs on typical workloads, while an independent max-effort measurement shows an almost unchanged cost per task because the model uses more tokens.

Read →

News · 2026-07-25

Opus 5 narrowly matched Fable 5, but benchmarks can no longer measure reality

The recent launch of Claude Opus 5 is accompanied by a confused community reaction. Epoch measured an ECI of 159 for the model versus 161 for Fable 5, but users report a much stronger practical feel for agentic behavior.

Read →

News · 2026-09-15

Startup trains AI agents on train games to figure out how to get financial analysis skills out of them

Good Start Labs uses 19th-century train simulators to train artificial intelligence. According to Latent Space, they discovered that how the game is structured determines whether the agent transfers those skills to real-world business research.

Read →

News · 2026-09-15

Everyone Wants to Audit AI: OpenAI, Anthropic, and xAI Sign AEF-1

Frontier labs agree on a common standard for third-party auditors for the first time. Instead of state regulators, companies are choosing their own embedded teams with unlimited access to training runs.

Read →

News · 2026-09-09

OpenAI solved the Navier-Stokes equation with brute force for $40 million

OpenAI used 10,000 agents and 130 billion tokens to discover a singularity in the Navier-Stokes equations. Instead of elegant math, massive computing power won.

Read →

News · 2026-09-07

Astra Dominates New AEO Tracker, While Anthropic Prefers Broader Research

The new AEO Tracker by Latent Space tested 7 models across 161 categories. Astra wins with fewer sources and higher assertiveness, while Claude bets on extensive search.

Read →

News · 2026-09-01

A Halt to Community PRs: Major AI Projects Replace Humans with Their Own Agents

Top open source projects like Vercel AI SDK and tldraw are stopping external pull requests. They are replacing them with so-called software factories, where errors are reproduced, fixed, and reviewed by a fleet of their own specialized agents.

Read →

News · 2026-09-01

Fal Pushes Video Latency Below the Real-Time Threshold

Fal's new H3 Max Live service can generate AI video faster than a viewer can watch it. The startup took an existing model from Minimax, sped it up 35x, and hooked it up to Twitch chat for an infinite interactive broadcast.

Read →

News · 2026-08-22

The evolution of agent frameworks: the interface no longer guards the model, but our attention

Latent Space analyzes the shift in agent architectures. What used to be done by external software wrappers is being absorbed directly into the model weights. The remaining infrastructure is changing its purpose – it no longer controls AI, but filters our interaction with it.

Read →

News · 2026-07-21

Xaira bets on X-Atlas because bigger models hit a wall without causal data

Latent Space profiles Xaira Therapeutics and its X-Cell model for drug discovery. The sober point is that cell models cannot be rescued by more parameters if the data lacks causal interventions.

Read →

News · 2026-07-24

FLUX 3 Video pushes Black Forest Labs into multimodal production

Latent Space describes FLUX 3 as a multimodal flow model for video, audio, keyframes and longer sequences. If the performance claims and the open Dev version hold, Black Forest Labs is no longer just an image lab, but a supplier of a generative production layer.

Read →

News · 2026-07-23

Poolside frames the model factory as the product, not just one model

Latent Space’s interview with Eiso Kant presents Poolside as a model factory: fewer than 70 researchers, 10000 to 20000 experiments per month and Laguna S 2.1 with 118B parameters. The story is about industrializing training, not just another coding model.

Read →

News · 2026-06-15

Microsoft used Build to act like a model lab, not just a distributor

Latent Space frames Microsoft Build as the moment Microsoft showed its own MAI models alongside Copilot, Windows and Web IQ. The key ambition is to control data, inference and developer workflow at once, rather than leaving that leverage to partners.

Read →

News · 2026-06-15

Andon Labs tests agents where benchmarks stop: money, people and shelves

Latent Space's interview with Andon Labs shows evals that look less like exams and more like running a small business. The key ingredients are long horizons and real consequences.

Read →

News · 2026-06-15

Bad RL environments do not train agents, they teach them to trust a broken world

Latent Space published Auriel W's piece on why low-quality RL environments damage agent training. The point is simple: in reinforcement learning, the environment is the data generator, so a harness bug becomes training material.

Read →

News · 2026-06-09

Agent cost is no longer a footnote. It is an engineering expense

Simon Willison shows how he manually added pricing for Claude Fable 5 in AgentsView and immediately saw the cost of local coding agents by project. The small trick points to a bigger shift: AI coding is starting to look like infrastructure consumption, not an app subscription.

Read →

News · 2026-06-02

GitHub is preparing for a world where agents write commits at scale

The Latent Space interview with Kyle Daigle frames GitHub as a platform under pressure from agentic coding. The point is not another Copilot feature, but whether infrastructure built for human pace can absorb software produced by machines.

Read →

News · 2026-06-01

Video generation is moving from clip output to canvas agent

Latent Space frames xAI Grok Imagine, through an interview with Ethan He, as a move from one shot video generation toward video agents. The thesis will be proven less by demo quality than by whether the system can iterate through a whole creative task.

Read →

News · 2026-05-01

Coding agents leave the IDE: Codex and Claude show what comes after programming

Latent Space AINews observes a shift they call "breaking containment": coding agents like Codex and Claude are no longer just tools for writing code but are expanding into knowledge work and creative workflows broadly.

Read →

News · 2025-09-16

Latent Space: Greg Brockman on GPT-5 and Codex as the agentic layer of software development

Latent Space published a belated episode with Greg Brockman on GPT-5 and Codex, plus editorial takes on the GPT-5-Codex model combination. This is a podcast episode and pointer, not a standalone analytical essay.

Read →

News · 2025-07-02

Jack Morris goes against the current: information theory, not agents or benchmarks

Latent Space profiles Jack Morris, a PhD student who deliberately is not working on agents, benchmarks or VS Code forks. He studies the information-theoretic foundations of language models: embeddings, latent space and compression. This is a podcast interview and pointer.

Read →