Behind the Craft · Sunday, August 2, 2026
Karan Malhotra highlighted that in-context learning (ICL) is the most powerful tool for AI models, more so than fine-tuning. He explained that by providing examples and saving desired behaviors, users can achieve test-time reinforcement learning within Hermes. This allows the model's context to become so overwhelming that it minimizes differences between various underlying models like Claude or GPT.
“The most powerful thing for a model is in context learning. ICL is more powerful than everything else, fine tuning, whatever.”
“You're doing a sort of test time reinforcement learning. You're doing a sort of test time improvement. And that test time improvement that stays only in the harness of memory, skill, increases, memory, skills, increases, the efficiency of a skill, self-improvement loop, the janitor maintenance inside of the harness.”
“When I switch from, uh, Claude to Chat GPT on the website, I get two totally different behaviors. When I switch inside of Hermes that has this very particular to me context, I barely will notice the difference in what I'm talking about because the context is so overwhelming to the model.”