The AI Daily Brief · Wednesday, September 9, 2026
Google has launched the Gemini 3.8 Flash model, a rapid iteration on their Flash series, emphasizing improved performance on complex tasks through iterative tool calls and reasoning. While competitive on some benchmarks, its performance on others, particularly Terminal Bench 4.0, lagged behind expectations, suggesting generalization challenges.
“The central claim from Google around 3.8 Flash is that it will work harder than 3.7. It's trained to call tools iteratively and perform more reasoning steps on complex tasks, yielding much better results.”
“On the benchmarks, the model looks solid, if a little spiky. It scored 73.7% on coding benchmark Deep Sweet, just a hair shy of Opus 5's score of 74%, and outperforming GPT-5.6 Soul by 1%.”
“Ultimately, Gemini 3.8 Flash remains, what's the most expensive and most capable models we've seen yet. Despite being benchmarked against models that cost significantly more.”