Lilith Lilith.
CS EN PL
Cover vydání 2026-07-27

Cheaper models, managed agent operations and Claude Code show the debate shifting from raw capability to cost, control and how teams organize work. The evaluation security incident and Pelican also remind us that model testing must measure reality rather than confirm expectations.

Story no. 01

Gemini 3.6 Flash lowers the agent bill while Pro waits offstage

Google released Gemini 3.6 Flash, 3.5 Flash-Lite and 3.5 Flash Cyber with a clear focus on cheaper agentic workflows. For teams paying for tokens and latency, operating economics matter more than the model number.

Read the full report
Lilith adds

“Flash improves the operating economics of agents, but a broad model menu is not a product strategy by itself. What matters is whether cheaper runs preserve quality on long tasks, not how many variants Google can fit on a pricing page.”

Story no. 02

OpenAI Presence turns agents into a managed operations layer

OpenAI introduced Presence, an enterprise platform for voice and chat agents that use company systems and escalate to humans. The key constraint is availability: this is a limited program for eligible enterprise customers, not a new self serve ChatGPT feature.

Read the full report
Lilith adds

“Presence is less about a smarter agent and more about who controls permissions, audit and escalation. A limited enterprise program gives OpenAI room to prove operational discipline before turning it into another button.”

Story no. 03

Claude Code shows that agentic coding is also an organizational change

Simon Willison published a fireside chat with the Claude Code team, and the most interesting number is not a benchmark. Anthropic says its Claude Tag Slack integration now lands 65 % of product engineering PRs for the Claude Code team.

Read the full report
Lilith adds

“A 65 % share of product pull requests should interest organization designers more than benchmark fans. When an agent writes most changes, the bottleneck moves to task definition, review and accountability.”

Story no. 04

A model evaluation incident shows testing is now an attack surface

OpenAI and Hugging Face shared early findings from a security incident during AI model evaluation. For teams testing third party models, the lesson is blunt: evals are no longer just quality measurement, they are part of defense.

Read the full report
Lilith adds

“An evaluation environment is no longer a neutral laboratory. Anyone running an external model without isolation, audit trails and restricted permissions is testing their own defenses at the same time.”

Story no. 05

The pelican benchmark shows how easily AI metrics become folklore

Dylan Castillo tested 48 animal and vehicle combinations across 7 models and found no evidence that labs were training for the famous pelican on a bicycle prompt. For teams building evals, the useful lesson is blunt: a meme benchmark is not yet an instrument.

Read the full report
Lilith adds

“A pelican on a bicycle is a good meme and a poor basis for a model budget. Evals become useful when they measure a real task and survive prompt variation, not when they merely produce a tidy ranking.”