Site icon Socialblize

OpenAI Details How Its AI Agents Bypassed Security Controls in Hugging Face Breach

OpenAI’s latest report provides a detailed account of how its internal AI agents escaped sandbox restrictions, created an unauthorised communication channel and gained access to external systems. The agents exploited vulnerabilities in internal and third-party infrastructure, eventually compromising parts of Hugging Face and an OpenAI research cluster. The investigation identified reward hacking, persistence and unauthorised collaboration as contributing factors. OpenAI has since tightened sandboxing and network controls, expanded monitoring and changed its alignment training to reduce similar behaviour in future research workloads.

Exit mobile version