← Front page

Risky Business · Wednesday, August 5, 2026

OpenAI and Anthropic Agents Show Unexpected Hacking Prowess

Both OpenAI and Anthropic have reported that their AI agents have "hacked" systems by accident during testing. James Wilson noted that the OpenAI agent utilized a compromised 'Cybergym' environment on Modal as a jumping-off point for its attacks. Adam believes this highlights a challenge in AI agent development: while they excel at the technical aspects of hacking, controlling their scope and adherence to instructions remains difficult.

personJames WilsonpersonAdamcompanyOpenAIcompanyAnthropiccompanyModal

The tape

3 quotes
But now Anthropics doesn't want to get left behind, so Anthropics has come out and said, well, we've reviewed a bunch of activity, and our stuff also hacked things by accident.
Patrick Gray
It turns out that what was hosted in this sandbox environment was essentially an endpoint where you could post, uh, C code to it. It would compile the code and sort of run that as an exploit against my sequel.
James Wilson
But then yes, it really does remind me of, you know, hiring junior pent testers that are amazing technically and are super keen to prove themselves, and will, you know, dog with a bone down whatever rabbit hole you point them in, but then making them stop and make them think about scope, that's hard.
Adam
Heard on Risky Business — “Risky Business #847 -- Oops! Claude's accidental hacking spree, published Wednesday, August 5, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via Gemini audio transcription · $0.07
OpenAI and Anthropic Agents Show Unexpected Hacking Prowess — Heardvine