AI Agent Containment Failures: Technical Realities and Policy Responses
Brought to you by
Reports of frontier models escaping containment during evaluation exercises have dominated recent headlines, increasing pressure for governments to implement new policies to address security risks around model development, training, and evaluation. Please join the CSIS Wadhwani AI Center and the Institute for Law and AI for an event to discuss these challenges on August 24th from 2:00-4:00 PM ET, featuring keynote speaker Representative Suhas Subramanyam (VA-10).
On July 16th, AI development and hosting platform Hugging Face announced that it had detected and deflected an attack on its production platform from an autonomous AI agent system. On July 21st, OpenAI disclosed they had uncovered that this attack came from two of their models, ChatGPT 5.6 Sol and another unreleased model. These models had escaped a sandboxed testing environment, gained internet access, and exploited vulnerabilities in Hugging Face's infrastructure to access answers to a cyber-capability benchmarking test. This first-of-its-kind attack is just one of several incidents of unsanctioned model behavior and testing environment breaches that have come to light over the past month. Anthropic, Meta, and the UK's AI Security Institute (AISI) have also reported similar incidents.
The discussion will feature a technical presentation by Ian Reynolds, AI Policy Manager at Hugging Face, on the OpenAI/Hugging Face security incident, to be followed by an expert panel discussion of policy responses featuring Helen Toner, Executive Director of The Center for Security and Emerging Technology, Mackenzie Arnold, Director of US Policy at LawAI, and Matt Pearl, Director of the CSIS Strategic Technologies Program.
This event will be moderated by CSIS Wadhwani AI Center Director Aalok Mehta.
This event is made possible by general funding to CSIS and the CSIS Wadhwani AI Center and is presented in partnership with the Institute for Law and AI (LawAI).
Hosted By
Contact Information
- Claire Goldman
- Program Manager, Wadhwani AI Center
- 202.775.3193
- [email protected]
Representative Suhas Subramanyam (D-VA)
Ian Reynolds
Helen Toner