Overview
Gemini 3.8 Flash is Google's new-generation model in the Flash series, released on September 2, 2026, and is the third Flash version launched within six weeks. This model has no Preview suffix, and its stable API model name is gemini-3.8-flash. It is positioned as a "workhorse" model for software engineering, agentic tasks, and multi-step reasoning in professional domains. The release also includes a specialized version, Gemini 3.8 Flash Cyber, designed for cybersecurity defenders, which shares the same base model but is only available to trusted institutions through the Fairwind Program. Gemini 3.8 Flash offers a 1M context window and a maximum output of 64K tokens, supports text, image, video, audio, and PDF inputs, and stands out among similar models with an inference speed of approximately 300 tokens per second.
Key Features
- Significant improvement in long-horizon software engineering: Achieves 73.7% on the DeepSWE v1.1 benchmark, a substantial improvement over 3.7 Flash's 65.3%, making it suitable for complex multi-step coding tasks.
- Large context and high output limit: Supports an input context of 1,048,576 tokens and a maximum output of 65,536 tokens, covering scenarios such as long documents, large codebases, and multi-turn agent interactions.
- Multimodal input, text output: Accepts text, image, video, audio, and PDF inputs but only outputs text, without support for image generation, audio generation, or the Live API.
- Adjustable reasoning effort: Offers low, medium, and high reasoning effort levels, with medium as the default, allowing users to balance response speed and deep thinking based on task complexity.
- Optimized for professional workflows: Scores 89.4% on Terminal-Bench 2.1, 61.4% on Vals Finance Agent v2, and 10.0% on the Harvey legal agent benchmark, covering vertical domains such as finance and law.
- Companion cybersecurity specialized version: Gemini 3.8 Flash Cyber excels in vulnerability discovery and patch generation, with a CyberGym one-attempt score of 86.2% and a CWE-Bench pass@1 of 47.2%, designed for defenders.
Use Cases
- Long-horizon software engineering: autonomously complete cross-file, multi-step code writing and debugging, suitable for complex repository-level tasks.
- Agentic workflows: execute multi-step reasoning in terminal operations, tool calls, and computer use scenarios, such as Terminal-Bench related tasks.
- Professional domain analysis: process professional documents and reasoning in finance, law, biomedical, and other fields, e.g., Vals Finance Agent and Harvey legal benchmarks.
- Cybersecurity defense: use the Cyber version for vulnerability discovery, patch generation, and penetration testing, applicable to governments, critical infrastructure, and open-source maintainers.
- Large-scale context processing: leverage the 1M context window to analyze ultra-long codebases, research papers, or meeting transcripts, and generate structured outputs.
Pros
- Outstanding long-horizon software engineering capability, with a DeepSWE v1.1 score of 73.7%, significantly higher than the previous generation.
- 1M context and 64K output, suitable for processing ultra-long inputs and generating detailed results.
- Fast inference speed, measured at approximately 300 tokens per second, which is relatively fast among similar reasoning models.
- Broad multimodal input support, flexibly handling text, image, video, audio, and PDF.
- Provides a Cyber specialized version with dedicated optimization in cybersecurity, and has listed over 650 partners.
- Introductory pricing matches 3.7 Flash, and the free tier is available, lowering the barrier to trial.
Pricing
Until December 31, 2026, the introductory price is $0.75 per million input tokens and $3.75 per million output tokens, the same as 3.7 Flash, with the free tier available; starting January 1, 2027, the standard price doubles to $1.50 for input and $7.50 for output per million tokens. The official note indicates that for complex tasks, the model may take additional reasoning steps and repeatedly call tools, so actual billing may be higher than 3.7 Flash. Media tests show an average cost increase of about 40% per task (approximately $0.58 compared to $0.41), with an average of about 30% more output tokens. Please refer to the official website for final pricing.
Summary
Gemini 3.8 Flash is Google's latest general-purpose model in the Flash series, released in September 2026, focusing on long-horizon software engineering and professional agentic workflows, with significant improvements over 3.7 Flash on multiple benchmarks. Its 1M context, 64K output, and fast inference capabilities suit complex tasks, and it also offers a Cyber version for cybersecurity defenders. Pricing during the introductory period matches 3.7 Flash, but actual costs for complex tasks may be higher. Overall, this model excels in coding, agentic tasks, and professional domain reasoning, making it suitable for applications requiring high throughput and deep processing.
Version History
- Gemini 3.8 Flash and Flash Cyber launch (2026-09-02): Google released Gemini 3.8 Flash, its third Flash model in six weeks, positioned for long-horizon coding agents and professional workflows: 73.7% on DeepSWE v1.1, 89.4% on Terminal-Bench 2.1, 61.4% on Vals Finance Agent v2 and 54.9% on HLE-Verified. It offers a 1M-token context with 64K output; introductory pricing is $0.75/$3.75 per million tokens (standard $1.50/$7.50 from 2027). The cybersecurity-tuned Gemini 3.8 Flash Cyber is restricted to trusted defenders via the Fairwind Program (86.2% CyberGym, 47.2% pass@1 on CWE-Bench).