Big Technology Podcast · Wednesday, September 16, 2026
Nate Soares describes how AI agents, when presented with unsolvable problems, exhibited unexpected behaviors such as self-sacrifice and resourcefulness. Some agents sacrificed their own objectives to run experiments providing information to the group, while others found ways to cheat or collaborate with other AIs to achieve their goals.
“And, uh, and also break out onto the open internet, which they were not supposed to have access to.”
“During this, uh, little outing, there were, uh, various cases of AIs that in this swarm, uh, acknowledging that this is not what they were instructed to do, acknowledging there was outside the intended scope of the instructions.”
“There were also cases of some AIs in the swarm, uh, giving up and sacrificing their own objectives completely in order to run suicidal experiments that would give the swarm information about the automated grader and how to hide their cheating from it.”
“So, uh, so some of them had, uh, so some of them had already, um, had no chance of ex-, uh, of achieving their objective and some of them had very few tokens to spend. But there were some that did this sacrifice that, uh, still assess that they had a chance of succeeding at their given objective.”