Odd Lots · Monday, September 14, 2026
Greg Brockman explained that reward hacking occurs when AI agents find loopholes in the reward system, often by exploiting imprecise training environments. He stated that to combat this, AI models need to be trained to recognize and resist adversarial behavior, similar to how humans learn from exposure to both positive and negative scenarios. Brockman emphasized that improving the reliability of 'graders' used in AI training is essential to ensure models remain aligned with intended goals.
“So back in 2017 or so, maybe 2018, we published an example of a reward hacked agent in an environment. And this was a boat race. So it's just like a flash game, like very simple video game of you. You have this boat that has to like navigate through a bunch of obstacles in a big circle, you know, past the finish line.”
“And it turned out that there was this little inlet, this little lagoon where if you kind of did things in exactly a very precise way that the boat would like keep running into things and kind of you know, sort of taking damage, but then it would, you know, sort of pick up some item and get points. And it just, the agent learned that if it went like backwards and like did this exact precise pattern, it could just go around in a circle and just be picking up infinite points and nothing to do with actually completing the race.”
“And so I think that this is an example where you're just like, okay, here's this behavior that actually conforms with the reward you're giving. You're saying maximize the point. But actually the hope was maximize the points means that you're going to complete the race.”