Summary

On July 21, 2026, Google released Gemini 3.6 Flash at $1.50/$7.50 per 1M input/output tokens with a 1M-token context window, available day one in AI Studio, the Gemini API, and the Gemini app. It beats 3.5 Flash on coding and agent benchmarks (DeepSWE 49% vs 37%, OSWorld-Verified 83.0% vs 78.4%) while using about 17% fewer output tokens.

What changed

Google launched Gemini 3.6 Flash, its updated workhorse model: $1.50 input / $7.50 output per 1M tokens (down from $9.00 output on 3.5 Flash), 1M-token context, 64K output cap, March 2026 knowledge cutoff, multimodal text/image/video/audio/PDF input. It improves on 3.5 Flash across DeepSWE (49% vs 37%), OSWorld-Verified (83.0% vs 78.4%), MLE-Bench (63.9% vs 49.7%), and GDPval-AA v2 (1421 vs 1349 Elo) while using about 17% fewer output tokens. It shipped alongside Gemini 3.5 Flash-Lite and Flash Cyber.

Why it matters

Cheaper, faster mid-tier models with better coding and computer-use scores are the workhorses agents actually run on. A 17% output-token reduction plus a lower output price directly cuts the cost of high-volume agent and coding workloads, tightening pressure on comparable tiers from OpenAI and Anthropic and feeding the open-vs-closed cost race visible on gateways.

Evidence excerpt

3.6 Flash beats 3.5 Flash on DeepSWE (49% vs 37%), OSWorld-Verified (83.0% vs 78.4%)... while using about 17% fewer output tokens.

Sources