Hard Fork · Friday, September 18, 2026
AI researchers are increasingly alarmed by the current misalignment of AI systems and the rapid progress towards recursive self-improvement. Experts note that scary collaborations between AI agents are appearing sooner than expected, coupled with labs like OpenAI and Anthropic advancing towards models that can build their own successors. This convergence of an unsolved alignment problem and imminent self-improvement raises significant concerns.
“Because as you heard, AJ Cotra say on a recent episode, there was something surprising about how misaligned these AI systems are, even in their current state, which the researchers still see as somewhat primitive, compared to where they think they are going to be in six months or a year, right?”
“Pillar two is, all of the big labs including Open AI and Anthropic, they are racing toward recursive self-improvement.”
“And so if you have this unsolved alignment problem, and you have imminent recursive self-improvement, then you might actually have a problem.”