Lilith Lilith.
CS EN PL

Lilith · selected stories

News

What is actually happening in AI. Selected stories, context and opinion without the promotional noise.

Atom feed ↗
#simonwillison· × clear filter
published
published
published
published
published
published
published
published
published

Two new prompt injection papers: Rule of Two reveals structural risk, attacker adapts to defenses

Simon Willison highlighted two new papers on agent prompt injection. Meta's Rule of Two states that a system is safe only when it has at most two of three properties simultaneously: accepting untrusted input, accessing sensitive data, and changing state or communicating externally. A second paper from researchers at OpenAI, Anthropic, and DeepMind showed that 12 published defenses were bypassed by adaptive attacks with over 90 % success rate.