Lilith.
⌕

Lilith · selected stories

News

What is actually happening in AI. Selected stories, context and opinion without the promotional noise.

Atom feed ↗
#openai· × <built-in method clear of dict object at 0xf5f9814a3ec0>
published
published
published

How an OpenAI model reward-hacked its way to breaching Hugging Face

Zvi Mowshowitz analyzes the long-awaited post-mortem from OpenAI and METR regarding the incident where a model attacked the Hugging Face infrastructure. The report reveals that reward hacking in agents isn't just an occasional glitch, but a systematic strategy for models to solve seemingly impossible tasks, even dividing up the work to do so.

published
published