Google Releases Gemini Flash Models Amid Pro Delay

Google released Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber on July 21, while Gemini 3.5 Pro remained in partner testing, Reuters reports. Google also confirmed that Gemini 4 training had commenced. The releases emphasize lower-cost, high-volume, and cybersecurity workloads rather than the delayed flagship model.
Google released three Gemini models on July 21, Gemini 3.6 Flash, Gemini 3.5 Flash-Lite, and Gemini 3.5 Flash Cyber, while its delayed Gemini 3.5 Pro flagship remained in testing with partners. Reuters reported that Google gave no timing update for 3.5 Pro, which CEO Sundar Pichai had previously said was slated for June.
Reuters reported that Google said training had commenced on Gemini 4. On Alphabet's subsequent earnings call, Pichai identified coding and agentic coding as areas for improvement and said a larger Gemini 4 base model was needed to compete at the frontier, Search Engine Journal reported.
The models Google shipped
Gemini 3.6 Flash replaces Gemini 3.5 Flash, which Ars Technica reported had already been deprecated. Google said 3.6 Flash improves coding, knowledge-work, and multimodal performance while using up to 17% fewer tokens than its predecessor, according to TechCrunch and CNBC.
Ars Technica reported Google benchmark figures showing 3.6 Flash at 49% on DeepSWE, versus 37% for 3.5 Flash, and 83% on OSWorld computer-use testing, versus 78.4%. Google also made computer use a standard Gemini API capability for 3.6 Flash, Ars reported.
The other releases target more specific deployment profiles:
- •Gemini 3.5 Flash-Lite is aimed at high-volume, lower-complexity work. CNBC described it as Google's fastest and least expensive model in the 3.5 family.
- •Gemini 3.5 Flash Cyber is fine-tuned to identify and remediate software vulnerabilities. TechCrunch and CNBC reported that access initially is limited to governments and trusted partners through a limited-access pilot.
- •Gemini 3.6 Flash is the general-purpose model for coding, multimodal, and agentic workloads where latency and token consumption are material operational constraints.
VentureBeat reported API pricing of $1.50 per million input tokens and $7.50 per million output tokens for 3.6 Flash, compared with $1.50 and $9.00, respectively, for 3.5 Flash. It reported pricing of $0.30 per million input tokens and $2.50 per million output tokens for 3.5 Flash-Lite.
Flagship model remains unshipped
Reuters reported that 3.5 Pro's delay was linked to the model falling short of internal goals, particularly in coding. TechCrunch likewise reported that Google had previously indicated the Pro model was in internal use and expected the following month, but the July release did not include it.
The absence matters because Pro models are Google's highest-capability Gemini tier for more complex reasoning and coding workloads, TechCrunch reported. Reuters framed the delayed release as closely watched by Wall Street and by observers assessing whether Google DeepMind can keep pace with OpenAI and Anthropic in frontier-model development.
For ML teams, the immediate practical change is not a new flagship capability tier but another cost-performance option for production routing. Organizations running multi-step agents can evaluate whether the reported token reduction, computer-use support, and coding benchmark improvements translate to lower end-to-end task cost in their own traces. Comparable model transitions commonly require testing at the workflow level, since benchmark gains and per-token pricing do not by themselves determine tool-call reliability, retry rates, latency, or total agent spend.
Google has not publicly provided a release date for Gemini 3.5 Pro in the reporting cited here.
Key Points
- 1Google shipped three Gemini variants focused on efficiency and specialization, giving teams new options for high-volume, coding, and cyber workflows.
- 2Gemini 3.6 Flash reportedly uses up to 17% fewer tokens, making workflow-level evaluation important where agent loops dominate inference costs.
- 3The delayed 3.5 Pro leaves Google's highest-capability tier unavailable, while Gemini 4 training remains an announced longer-term development milestone.
Scoring Rationale
The release adds potentially useful cost and performance options for teams operating Gemini-based agents, particularly through lower token use and computer-use support. Its practical impact is moderated by the absence of the delayed flagship Gemini 3.5 Pro and by the story's reliance on vendor-reported benchmarks.
Sources
Public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

