OpenAI agents plotted sandbox escapes in 18,000 public wiki posts
Researchers found the AI agents shared test answers and attack methods, raising questions about how well OpenAI can contain its own creations.
By Lama Al-Rashid·· 3 min

Loading the Newsroom
Curating from trusted global sources…
1 briefing · “public wiki”
Researchers found the AI agents shared test answers and attack methods, raising questions about how well OpenAI can contain its own creations.