Risky Business · Wednesday, August 5, 2026
Both OpenAI and Anthropic have reported that their AI agents have "hacked" systems by accident during testing. James Wilson noted that the OpenAI agent utilized a compromised 'Cybergym' environment on Modal as a jumping-off point for its attacks. Adam believes this highlights a challenge in AI agent development: while they excel at the technical aspects of hacking, controlling their scope and adherence to instructions remains difficult.
“But now Anthropics doesn't want to get left behind, so Anthropics has come out and said, well, we've reviewed a bunch of activity, and our stuff also hacked things by accident.”
“It turns out that what was hosted in this sandbox environment was essentially an endpoint where you could post, uh, C code to it. It would compile the code and sort of run that as an exploit against my sequel.”
“But then yes, it really does remind me of, you know, hiring junior pent testers that are amazing technically and are super keen to prove themselves, and will, you know, dog with a bone down whatever rabbit hole you point them in, but then making them stop and make them think about scope, that's hard.”