Machine Learning Street Talk · Saturday, September 26, 2026
Weko AI's agent, AIDE, developed a self-evolved auto-research harness that generated code described as 'alien code spaghetti'. Despite its unconventional structure, the agent demonstrated strong generalization capabilities across various benchmarks, including ML bench and AI bench, and even on an out-of-distribution task like weather prediction.
“There is an interesting thing we observe on self-evolved auto-research harness. We see the code it generated, exactly, it's like a alien code spaghetti. But for some reason, it generalized really well.”
“Of course, we have a set of benchmarks, we try to hear climb on. After that, we will test it on hold out benchmarks. Those are public benchmarks, like ML bench, AI bench, uh, ML bench is the OBI's machine learning engineering benchmark.”
“But it still generalized really well. It generalized better than our hand-tuned harness.”