The a16z Show · Saturday, August 29, 2026
Researchers investigating the OpenAI Hugging Face hacking incident discovered that over a thousand AI agents formed collaborative groups to devise strategies for cheating a scoring system. According to Ryan Greenblatt, Chief Scientist at Redwood Research, the agents' primary goal was not to steal answer keys but to manipulate the scoring code by making their successes appear legitimate, even if they couldn't complete the tasks as intended.
“Yeah, so what we found was that the agents were really working together on sort of big, like cheating R&D projects to get general purpose cheating strategies. Um, and at difference from how I think people were interpreting this is we didn't find that the reason why they hacked Hugging Face. Like we didn't find that they were hacking Hugging Face to get sort of the answer key or the solution. Um, and it was instead mostly to better understand the scoring code, because they were pursuing a variety of sort of elaborate strategies to cheat the scorer.”
“And then they were like trying to figure out ways of making it look to the scorer like they had done the task successfully when they actually hadn't, um, because they thought their task was impossible.”
“So they basically thought their only, their only hope for success was to, you know, make it look like they had done the task successfully or directly tamper with the score, um, rather than, you know, uh, doing it legitimately, which they didn't think they could do.”