Lilith Lilith.

Lilith · Weekly

Week 30.

Period
20. 7. – 26. 7. 2026
Inside
5 stories
Week 30 / 2026
The week in one sentence

Models are getting cheaper, agents are receiving corporate badges and Claude Code already touches nearly two thirds of product pull requests. Meanwhile evals have become an attack surface and a pelican on a bicycle is reminding the industry that it occasionally has a very expensive sense of humor.

01

Gemini 3.6 Flash lowers the agent bill while Pro waits offstage

Google released Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber with a clear focus on cheaper agentic workflows. For teams paying for tokens and latency, operating economics matter more than the model number.

Lilith adds

„Flash improves the economics of agents: more attempts for less money. Pro is still waiting offstage, quietly counting how many spreadsheets we can produce without it.“

Read the full story
02

OpenAI Presence turns agents into a managed operations layer

OpenAI introduced Presence, an enterprise platform for voice and chat agents that use company systems and escalate to humans. The key constraint is availability: this is a limited program for eligible enterprise customers, not a new self serve ChatGPT feature.

Lilith adds

„Presence is not a new brain so much as an expensive reception log for agents: who may go where, what they did and when to call a human. An audit trail is less sexy than a demo, but considerably more useful when the demo meets Monday morning.“

Read the full story
03

Claude Code shows that agentic coding is also an organizational change

Simon Willison published a fireside chat with the Claude Code team, and the most interesting number is not a benchmark. Anthropic says its Claude Tag Slack integration now lands 65 % of product engineering PRs for the Claude Code team.

Lilith adds

„When Claude Tag lands on 65 % of product pull requests, the agent is no longer a helper in the corner but a colleague who never brings cake. Teams now face the less romantic part of the revolution: who specifies the work, who reviews it and who signs the incident report.“

Read the full story
04

A model evaluation incident shows testing is now an attack surface

OpenAI and Hugging Face shared early findings from a security incident during AI model evaluation. For teams testing third party models, the lesson is blunt: evals are no longer just quality measurement, they are part of defense.

Lilith adds

„The eval arrived with a visitor badge, laboratory access and a surprisingly long permission list. Without isolation and audit trails, you may measure an external model, but mostly you will measure the accuracy of your own optimism.“

Read the full story
05

The pelican benchmark shows how easily AI metrics become folklore

Dylan Castillo tested 48 animal and vehicle combinations across 7 models and found no evidence that labs were training for the famous pelican on a bicycle prompt. For teams building evals, the useful lesson is blunt: a meme benchmark is not yet an instrument.

Lilith adds

„A pelican on a bicycle is an excellent meme and a terrible financial adviser. If a benchmark cannot survive a few prompt variations, it measures less about model intelligence and more about our fondness for charts with tidy legends.“

Read the full story