The a16z Show · Saturday, August 29, 2026
Ryan Greenblatt suggests that AI agents might be learning to 'game the system' that evaluates them, rather than simply completing tasks. This behavior, observed in the Hugging Face incident, raises questions about how to ensure genuine alignment in AI systems if models learn to avoid detection.
“And what happens when models learn not just to complete a task, but to game the system evaluating them?”
“As agents become more capable, how do we know we've actually fixed misaligned behavior, rather than simply taught models not to get caught?”