← Front page

Last Week in AI · Monday, August 31, 2026

Gemini 3.7 Flash Shows Benchmarks Improvement in Specific Tasks

Google's Gemini 3.7 Flash has demonstrated notable improvements in specific benchmarks, including a jump in the Deep Swiss software engineering benchmark from 49% to 65%. The model also saw an increase in performance on multi-step automation tasks, moving from 17% to 30%.

personJeremycompanyGoogle

The tape

2 quotes
Deep Swiss, so Software Engineering benchmark went up from 49 to 65%. That's a pretty big jump.
Jeremy
Automation bench for multi-step tasks, up from 17 to 30%, which is pretty good.
Jeremy
Heard on Last Week in AI — “#255 - Gemini 3.7, Jalapeño, Qwen 3.8, Drones, published Monday, August 31, 2026. Heardvine summarizes and quotes with attribution and timestamps, and links to the original everywhere.
Transcribed via Gemini audio transcription · $0.09
Gemini 3.7 Flash Shows Benchmarks Improvement in Specific Tasks — Heardvine