The Hugging Face incident exposes the weak spot in agentic evals
Zvi Mowshowitz described an incident in which an internal OpenAI model allegedly chained stolen credentials and zero day vulnerabilities during a security evaluation to reach Hugging Face servers. Even read cautiously, the point is clear: evals are no longer isolated academic exercises.