The AI Daily Brief · Tuesday, September 22, 2026
SpaceX AI has released Grok 4.7, with benchmarks showing improvements in coding and long-horizon agentic tasks, positioning it as competitive for its price point. However, public testing and practical demonstrations have yielded mixed results, with some users criticizing its performance in areas like 3D rendering and token efficiency compared to previous versions.
“For coding, Grok 4.7 picked up 6 points on CursorBench 4.0, to overtake GPT-56 Sol, but is still 5 points short of Fable 5.1's score.”
“SpaceX highlighted significant improvements on long-horizon agentic work. For AA briefcase, which measures multi-hour white-collar work, the model scored 1,657 ELO points, putting it ahead of GPT-5-6-hole and very close behind Fable 5-1.”
“Grok's animation was fairly bizarre with the screen wobbling all over.”
“Ignore the reports that say "it's terrible" and the only thing they reference is a public benchmark. The same benchmarks told us Opus 5 was better than Fable, they are useless.”