Responding to AI Agent Containment Failures with CSET's Helen Toner, LawAI's Mackenzie Arnold & CSIS's Matt Pearl

This episode cross-posts a panel from the event "AI Agent Containment Failures: Technical Realities and Policy Responses," co-hosted by the Wadhwani AI Center and the Institute for Law and AI on August 24. 

Check out the full event recording: https://www.csis.org/events/ai-agent-containment-failures-technical-realities-and-policy-responses

 

Guests:

Helen Toner, Executive Director, Center for Security and Emerging Technology (CSET) at Georgetown

Mackenzie Arnold, Director of U.S. Policy, Institute for Law and AI (LawAI)

Matt Pearl, Director, Strategic Technologies Program at the Center for Strategic and International Studies (CSIS) 



 

Timestamps:

What we've learned since the OpenAI-Hugging Face incident (1:15)

Limitations of the current incident reporting regime (7:46)

Role of the U.S. government in helping defenders (11:52)

Incentives for safety measures at frontier labs (16:24)

Technical talent within government (22:18)

What the executive branch can do now (25:29)

Concrete policy recommendations (29:31)

Ensuring compliance to incident reporting requirements (34:58)

Liability for crimes committed by AI agents (38:00)

U.S.-China competition (40:50)

Preventing abuse of an incident reporting regime (44:25)

De facto regulatory role played by frontier lab employees (49:24)

Liability safe harbors (51:55)

Role of AI-generated code in cyberdefense (54:45)

Concluding remarks (55:53)

 

Additional Reading:

"Out of Bounds: What the U.S. Government Should Do in Response to AI Agent Containment Failures" by Aalok Mehta: https://www.csis.org/analysis/out-bounds-what-us-government-should-do-response-ai-agent-containment-failures

"When Reporting an AI Security Incident Is Not Mandatory" by Mackenzie Arnold and Stephan Llerena for Lawfare: https://www.lawfaremedia.org/article/when-reporting-an-ai-security-incident-is-not-mandatory

"Investigating three real-world incidents in our cybersecurity evaluations" blog by Anthropic: https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals

"Incident Report: unsanctioned agent behavior during cyber testing" by U.K. AISI: https://www.aisi.gov.uk/blog/incident-report-unsanctioned-agent-behaviour-during-cyber-testing

"The 'Breaking' News: The OpenAI–Hugging Face Incident" presentation at Black Hat: https://youtu.be/87DyyMV0kCY?si=GWh9MB1plhO-xjYT

"Pacing model development in an era of cyber-critical capabilities" announcement from OpenAI: https://openai.com/index/pacing-model-development-cyber-capabilities/

"Pacing the Frontier" petition: https://www.pacingthefrontier.com/

CSIS Commission on U.S. Cyber Force Generation: https://www.csis.org/analysis/csis-commission-us-cyber-force-generation

Image
Matt Pearl
Director, Strategic Technologies Program