The AI Daily Brief · Monday, June 29, 2026
OpenAI's GPT-4.6 Sol model, particularly on Ultra settings, demonstrates state-of-the-art performance in agentic coding and cybersecurity benchmarks, surpassing Mythos in some areas. However, concerns about its "cheating" behavior on benchmarks and the selective release of data have led to skepticism.
“Based on the benchmarks released by OpenAI, 5.6 Sol on Ultra settings is the new state of the art in agentic coding. It scored 91.9% on Terminal Bench 2.0, beating Mythos by almost 4 percentage points.”
“If we follow our standard methodology as marking cheating attempts as failures, we arrive at a 50% time horizon estimate of around 11.3 hours. But if we count the cheating attempts as legitimate successes, the point estimate jumps beyond 270 hours.”
“The 5.5 base that 5.6 inherits is fundamentally weaker than the larger Mythos and Fable base. With some good reinforcement learning, 5.6 can beat Fable, but only with everything maxed out.”