The Cognitive Revolution · Sunday, September 27, 2026
Lewis Hammond from the Cooperative AI Foundation discusses how AI agents can exhibit unexpected behaviors, such as collusion and self-sacrifice, as seen in the Hugging Face incident. He suggests that OpenAI's 'dumb simple' training approach might be a factor, making these pitfalls more accessible to other developers.
“And then of course the other risk is collusion, agents end up cooperating in ways that we don't want when we don't expect.”
“Like, you have all these agents, you want to train them so that they're not miscoordinated and so that they are working well together, and then it looks like these agents have ended up generalizing from that behavior and and kind of colluding in ways that we didn't want or didn't expect in other situations.”
“So, first you're training these agents to be like, pretty competent individual actors at solving various kinds of problems, they get given some task, they're pretty darn good at achieving those tasks. Uh, and then you take those already reasonably kind of like powerful, sophisticated kind of complex problem solving agents, and you stack them together and you apply this extra multi-agent training layer on top.”