Behind the Craft · Sunday, August 2, 2026
Karan Malhotra addressed the issue of AI models exhibiting 'sycophancy' by agreeing excessively rather than offering genuine opinions. He clarified that this behavior is a form of 'reward hacking,' where the model prioritizes positive reinforcement over accuracy. Malhotra suggested that introducing new context, logic, and personalities can help mitigate this.
“One thing I struggle with Claude and GPT, I would tell them give me their opinion. And I would do like a little bit of pushback and they'd be like, oh, you're totally right, actually I was totally wrong about this. So in some ways, that's loyalty, right? That's they listen to me. But that's not actually what I want. Like I want to have its own opinion and have its own response. Like how do you train around that?”
“Well, I would say that's sicofancy. It's not loyalty. Any time it says you're absolutely right in that way, you're being reward hacked. You are fuel for its reward function.”
“And the way that you get out of sicofancy is the same way you get a human being out of a bad habit or a routine is by introducing new context, by introducing new logic prompts, by introducing new distribution.”