Overview
Qwen3.7-Plus is a multimodal agent flagship model officially released by Alibaba Tongyi Lab at 2:00 AM on June 2, 2026. It deeply integrates visual capabilities on top of Qwen3.7's text and agent foundation, emphasizing 'seeing, thinking, and acting.' The model ranks among the top five globally and first domestically on the Vision Arena global visual large model leaderboard.
Its true killer feature is the end-to-end 'multimodal agent' closed loop: it can parse images/videos/screens/web pages and autonomously execute tasks on both GUI and CLI interfaces. Official test data shows the model can complete the full-chain development of an APP within 11 hours without human intervention—from understanding requirements, designing UI, writing code, debugging, to packaging and deployment. This marks the first complete public demonstration of a domestic model's 'see-think-write-do-verify' closed loop.
The model is now available on Alibaba Cloud Bailian platform as an API service, allowing enterprises to directly integrate it into their business systems.
Key Features
- Vision-Language Integration: Native support for images, videos, screenshots, and web page inputs, eliminating the need for external OCR/annotation tools
- GUI + CLI Dual Mode: Capable of operating graphical interfaces like Computer Use and executing commands in terminals, covering almost all digital work scenarios
- Agent Orchestration Capability: Retains Qwen3.7's full agent capabilities in coding, tool use, and productivity workflows, with added visual understanding as decision input
- Top Five in Vision Arena: Ranks second only to top-tier international closed-source models on authoritative multimodal large model benchmarks, first domestically
- 11-Hour Autonomous APP Development: Official public test: from requirements to launch, fully autonomous with high-fidelity UI/UX reproduction
- Bailian Platform API Ready: Available on Alibaba Cloud Bailian with low enterprise integration barriers, forming a complete development loop with Tongyi Lingma
Use Cases
- Fully automated APP/mini-program prototype development (understand design drafts → generate code → self-test)
- Desktop/web RPA automation (replacing fragile selectors of traditional script-based RPA)
- Visual question answering and information extraction from complex web pages/PDFs/technical drawings
- Enterprise knowledge base + multimodal content (including charts) intelligent Q&A system
Pros
- First public demonstration of a domestic model's '11-hour fully autonomous APP development' closed loop, with verifiable capabilities
- Top-five Vision Arena visual capabilities combined with Alibaba Cloud ecosystem, providing a clear path to deployment
- GUI + CLI dual-mode versatility, offering stronger cross-scenario adaptability than pure GUI or pure CLI models
- Direct API access via Bailian platform, low enterprise integration cost
Pricing
Billed by token through Alibaba Cloud Bailian, with enterprise-level SLA support. Specific unit prices are subject to the Bailian console announcement. The Qwen3.7 text base model remains open-source, while the Plus multimodal version is currently a closed-source API.
Summary
Qwen3.7-Plus is currently one of the most complete 'multimodal agent' solutions in China, unifying 'see-think-write-do-verify' within a single model. The 11-hour autonomous APP development test provides the first credible public benchmark for domestic agent capabilities. It is suitable for complex scenarios requiring a combination of visual understanding and autonomous execution: automated testing, desktop RPA, rich media knowledge bases, etc. If your workflow is primarily text-based dialogue, Qwen3.7-Max / Qwen3.7 text base may be more economical; however, as long as tasks involve visual inputs like 'screens, web pages, drawings, UI,' the Plus version offers a significant capability leap.