Overview
Gemini is a multimodal AI model series and conversational product (formerly Bard) launched by Google. As the core of the search giant's AI strategy, Gemini deeply integrates with Google Search, Gmail, Google Docs, YouTube, and other ecosystem services, making it one of the few AI assistants that seamlessly connects the entire Google suite.
**Gemini 3.5 series announced at Google I/O 2026 on 2026/5/20**: Currently officially available are **Gemini 3.5 Flash** (with significantly improved speed and agent capabilities, now the default search engine) and the full-modal **Gemini Omni** (gravity/kinetic simulation, conversational video editing). **⚠️ Note: The flagship model Gemini 3.5 Pro has not been officially released as of 2026/7/20** — originally planned for June, then rumored to be delayed to 7/17 for a "head-to-head" with the official version of DeepSeek V4, but according to a Bloomberg report on 7/17, Google employees revealed that Gemini 3.5 Pro's delivery has fallen months behind schedule due to **coding performance not meeting internal targets**. The company has only confirmed that it is "testing with partners" and has no new confirmed release date. This is one of the most talked-about "AI model delays" in the tech world in 2026, with the departure of several key researchers earlier also believed to be related to the delay pressure.
Key Features
- Deep Google Ecosystem Integration: Direct access to Gmail, Drive, Maps, and other services; after I/O 2026, supports "direct questioning of Gmail inbox" with fully natural language email search/summarization/context understanding
- 1M Token Ultra-Long Context: The industry's largest context window, with 3.1 Ultra further expanding to 2 million tokens of native context
- Native Multimodality: Supports text, image, audio, and video input; can analyze video content, understand screenshots, and listen to recordings
- Real-Time Web Search: Backed by Google Search, real-time information retrieval capability is top-tier among all AI assistants
- AI Studio Development Platform: Provides free API usage quotas and visual debugging tools, highly developer-friendly
- NotebookLM Knowledge Base: Build a personal knowledge base by uploading materials, supporting conversational retrieval and automatic podcast summary generation
- Gemini 3.5 Flash (Launched): High-speed, low-cost version, now the default Google Search AI engine, with significantly improved agent and coding capabilities
- Gemini Omni Full-Modal (Launched): Full-modal flagship announced at I/O 2026, seamlessly switching between text/image/audio/video/real-time interaction, with major upgrades in gravity simulation and kinetic computation
- ⚠️ Gemini 3.5 Pro (Not Yet Released, Delayed): Flagship model originally planned for June, then rumored for 7/17; according to Bloomberg, delayed due to coding performance not meeting internal targets, with no confirmed release date as of 7/20
Use Cases
- Heavy Google ecosystem users (daily Gmail, Drive, Docs users)
- Professionals who need to process ultra-long documents and video content analysis
- Researchers and information workers who need real-time information and web search
- Developers who want to experience powerful AI APIs for free (Google AI Studio offers generous free quotas)
- Creators who need full-modal input and output (Gemini Omni)
- Multilingual communication scenarios, with Gemini's broad multilingual capabilities
Pros
- Unmatched Google ecosystem integration: email summaries, document assistance, schedule management all in one
- 1M-2M ultra-long context: ability to process large files surpasses all competitors
- Top-tier web search: based on Google Search, with leading information timeliness and accuracy
- Free version is not weak: free access to Gemini Flash offers great value
- Native multimodal support: 3.5 Omni further breaks through in physical simulation and real-time interaction
- Complete product matrix after I/O 2026: flagship + lightweight + full-modal + Agent full coverage
Pricing
Gemini Basic is completely free, powered by the Gemini 3.5 Flash model. Gemini Advanced requires a Google One AI Premium subscription ($20/month), unlocking higher quotas for the 3.5 Flash / Omni series, 2TB cloud storage, NotebookLM premium quotas, and full Google ecosystem integration (**Note: 3.5 Pro has not been released yet and cannot be obtained via subscription**). For the API, Google AI Studio offers generous free usage quotas, while production environments are billed on a pay-as-you-go basis via Vertex AI, with the Omni model priced separately based on multimodal tokens.
Summary
After the I/O 2026 conference in May 2026, Gemini has launched the "Flash (lightweight) + Omni (full-modal) + Spark (Agent)" lineup, but the flagship **3.5 Pro continues to be delayed due to coding performance not meeting internal standards**, with no release date as of 7/20. This has prevented Gemini from competing head-to-head with the flagship models released in July, such as GPT-5.6, Grok 4.5, DeepSeek V4, and Kimi K3. For Google ecosystem users, Flash + Omni is already sufficient, making Gemini the most natural AI assistant choice; for developers, the free quotas of AI Studio are highly attractive; but if you are specifically waiting for 3.5 Pro for top-tier coding/reasoning capabilities, you need to keep an eye on official updates. It is recommended to use it alongside ChatGPT or Claude to complement each other's strengths.
Version History
- Google DeepMind 发布 Gemini 3.6 Flash、3.5 Flash-Lite 与 3.5 Fla (2026-07-21): Google DeepMind has launched three new models: Gemini 3.6 Flash, 3.5 Flash-Lite, and 3.5 Flash Cyber. Among them, Gemini 3.6 Flash is the latest flagship model, 3.5 Flash-Lite focuses on lower cost and higher efficiency, while 3.5 Flash Cyber is optimized for cybersecurity scenarios. All three models are available via API through the Google AI developer platform.
- Google 发布三款新模型:3.6 Flash、3.5 Flash-Lite 与 3.5 Flash Cyber (2026-07-21): Google has launched three new models aimed at improving performance, reducing latency, and lowering costs. Among them, 3.6 Flash reduces token usage by up to 65% on complex coding tasks, while 3.5 Flash-Lite achieves a speed of 350 output tokens per second. 3.6 Flash and 3.5 Flash-Lite are now available on the Gemini app, and 3.5 Pro has entered partner testing.
- OpenRouter 上线 Gemini 3.6 Flash 与 3.5 Flash-Lite (2026-07-21): Today launched on OpenRouter: Gemini 3.6 Flash and Gemini 3.5 Flash-Lite! Both are major updates to their model series, featuring high throughput (150+ tok/s), suitable for agent scenarios, from efficient token encoding and knowledge work to low-latency, high-concurrency sub-agents. Details below 🧵
- OpenRouter 推出 Classifiers 测试版:自动标记 AI 请求的用途与成本归属 (2026-07-24): OpenRouter has launched the beta version of Classifiers, allowing users to automatically tag each AI request with task type, department affiliation, compliance category, and other information through custom taxonomies (up to 8 dimensions). Classification runs asynchronously without increasing inference latency; it supports sampling rate control to manage costs, and recommends using Gemini 3.5 Flash Lite as the classification model. Tagging results are written to logs, and users can aggregate and analyze model usage distribution and cost flows by dimension in the Activity Explorer.
- Gemini 3.6 Flash Series (2026-07-21): 3.6 Flash reduces complex coding token usage by 65%; 3.5 Flash-Lite speed 350 tok/s; 3.5 Flash Cyber cybersecurity optimization
- Gemini 3.5 Pro Continues to Be Delayed (2026/07/17): Bloomberg report: Google employees reveal 3.5 Pro delayed for months due to coding performance not meeting internal targets; original June plan/rumored 7/17 release not realized; company only confirms testing, no new date; several key researchers previously left
- Gemini 3.5 Flash / Omni (Launched) (2026/05/20): Announced at I/O 2026. 3.5 Flash becomes default search engine, with significantly improved agent/coding capabilities; Omni is a full-modal breakthrough with major upgrades in gravity/kinetic simulation. Simultaneously launched Gmail natural language inbox (direct questioning of inbox), Pics smart image editing, 24/7 personal Agent Spark; flagship 3.5 Pro not released alongside Flash/Omni
- Gemini 3.1 Ultra (2026/02): 2 million token native context, built-in sandbox code execution tool, supports text/image/audio/video multimodality
- Gemini 3.1 Pro (2026/02): Comprehensive capabilities tied for top tier globally, leading in MMLU-Pro / GPQA and other benchmarks, 1M token context
- Gemini 2.x (2025): First to achieve 1M token long context, laying the foundation for multimodality
- Gemini 1.0 / Bard (2023-2024): Predecessor Bard upgraded to Gemini, competing with GPT-4, launching Google's AI strategy