Source: LLM Stats / GoogleAugust 20, 2026

Google Ships Gemini 3.7 Flash; Open-Source Model Wave Continues

View original source →

Google released Gemini 3.7 Flash this week, continuing its rapid Flash-tier model cadence. The company also updated Gemini 3.6 Flash with 17% fewer output tokens and reduced intermediate tool calls, directly lowering per-task API costs for high-volume agentic operations.

Key Points:

• Gemini 3.7 Flash released — following Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber in recent weeks • Flash-Lite now generates at 350 tokens per second — optimized for continuous document processing • NotebookLM rebranded to Gemini Notebook with new Canva, YouTube Music, and Instacart integrations

Open-Source Model Releases:

• DeepSeek released DeepSeek-V4-Pro-0813 and DeepSeek-V4-Flash-Vision-Exp • Alibaba's Qwen team released Qwen3.8-27B — a 27-billion parameter model for edge deployment • Zhipu AI released GLM-5.3 — focused on agentic tool invocation and multilingual task execution • xAI released Grok 4.6 for the Cursor IDE ecosystem

Open-weight models from Chinese labs now account for approximately 30% of global open AI inference volume, actively depressing pricing power of closed proprietary APIs. The capability gap with proprietary models is narrowing quarter over quarter.

Why It Matters: The sustained wave of capable open-source models from Asian labs is narrowing the capability gap with closed US frontier models — creating real procurement alternatives for enterprise teams.