Overview
Qwen3.8-Omni-Flash is a natively omnimodal model released by Alibaba's Qwen team on September 17, 2026, positioned to make audio and video agent tasks genuinely practical to run. It accepts text, images, audio and video in a single call, and its context window reaches 1M tokens, enough to hold an entire meeting recording or a long video for understanding instead of transcribing to text first.
Alibaba's own figures are that the average score across 29 benchmarks is more than 25% higher than the previous generation, Qwen3.5-Omni-Plus, while the hourly price for audio input falls by more than 98% and for audio-video input by more than 93%. The price movement matters more than the score, because the real bottleneck for omnimodal models has tended to be the cost of audio-video input rather than raw capability. Until that cost comes down, long-video understanding and real-time voice agents stay out of production reach.
Within the Qwen family, the Omni line covers multimodal input, a different job from the text-reasoning Qwen3.8-Max and Qwen3.8-Flash-Next models. Alibaba describes it as a model built for agentic delivery, meaning it does not stop at understanding audio and video but goes on to call tools and finish the task. For teams working with meeting recordings, customer service calls, surveillance footage or teaching material, this is the most directly relevant tier in the Qwen family today.
Key Features
- Native omnimodal input: Text, images, audio and video are handled in a single call rather than through separate pipelines for each modality, cutting down on engineering assembly work.
- 1M token context window: Large enough for a full long video or an extended meeting recording, enabling whole-document understanding and global summarisation instead of slice-by-slice processing.
- Agentic delivery for audio and video: Positioned to keep going after comprehension and call tools to finish the task, aimed at agents that need to watch, listen and act.
- Sharply lower audio-video cost: Alibaba says the hourly price of audio input drops by more than 98% and audio-video input by more than 93%, making budgets for long-material workloads easier to plan.
- Multimodal benchmark results: Alibaba reports an average score across 29 benchmarks more than 25% above Qwen3.5-Omni-Plus.
- Same-generation Qwen family fit: Shares the Qwen3.8 generation with the text-focused Qwen3.8-Max and Qwen3.8-Flash-Next models, so mixing and switching between them carries little adaptation cost.
Use Cases
- Structured notes from meetings and interviews: feed the audio in directly and get minutes, action items and decisions out
- Long video understanding: whole-video summaries, chapter splitting and key point extraction using the 1M token context
- Real-time voice and video agents: support, coaching and guiding scenarios where the model must watch, listen and act
- Teaching material processing: turning recorded lessons into searchable knowledge entries
- Multimodal asset review: bulk moderation and annotation assistance for audio and video material
Pros
- Native omnimodal design covers four input modalities through a single integration
- The 1M token context suits long audio and video material without pre-splitting and reassembling
- Audio and video input costs drop substantially, making long-material budgets easier to forecast
- Shares a generation with other Qwen models, keeping migration and mixed use inexpensive
- Alibaba's benchmark numbers show a clear multimodal improvement over the previous generation
Pricing
Called through the Alibaba Cloud Bailian platform and billed by input and output modality and by usage volume. Alibaba says audio input costs more than 98% less per hour than the previous generation and audio-video input more than 93% less; see the Bailian platform for current list prices.
Summary
Qwen3.8-Omni-Flash is the natively omnimodal model Qwen released on September 17, 2026, bringing text, image, audio and video input into one model with a 1M token context window. Alibaba's figures put the average score across 29 benchmarks more than 25% above the previous generation, alongside hourly price reductions of more than 98% for audio input and more than 93% for audio-video input. For teams handling long audio and video who want the model to keep going and call tools after it understands the material, it is the most directly relevant tier in the Qwen family.