ChengRang

Qwen3.8-Omni-Flash

AI Platforms Freemium
This page covers a version or sub-product of Qwen. View Qwen overview →

Natively omnimodal model released by Alibaba Qwen on 2026/9/17 with text, image, audio and video input and a 1M token context; Alibaba reports the average score across 29 benchmarks is more than 25% above Qwen3.5-Omni-Plus while hourly audio input pricing falls over 98% and audio-video input over 93%, aimed at audio and video agent delivery

OmnimodalAudio-VideoAgentQwen1M Context
Visit Qwen3.8-Omni-Flash

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

Qwen3.8-Omni-Flash is a natively omnimodal model released by Alibaba's Qwen team on September 17, 2026, positioned to make audio and video agent tasks genuinely practical to run. It accepts text, images, audio and video in a single call, and its context window reaches 1M tokens, enough to hold an entire meeting recording or a long video for understanding instead of transcribing to text first.

Alibaba's own figures are that the average score across 29 benchmarks is more than 25% higher than the previous generation, Qwen3.5-Omni-Plus, while the hourly price for audio input falls by more than 98% and for audio-video input by more than 93%. The price movement matters more than the score, because the real bottleneck for omnimodal models has tended to be the cost of audio-video input rather than raw capability. Until that cost comes down, long-video understanding and real-time voice agents stay out of production reach.

Within the Qwen family, the Omni line covers multimodal input, a different job from the text-reasoning Qwen3.8-Max and Qwen3.8-Flash-Next models. Alibaba describes it as a model built for agentic delivery, meaning it does not stop at understanding audio and video but goes on to call tools and finish the task. For teams working with meeting recordings, customer service calls, surveillance footage or teaching material, this is the most directly relevant tier in the Qwen family today.

Key Features

Use Cases

Pros

Pricing

Called through the Alibaba Cloud Bailian platform and billed by input and output modality and by usage volume. Alibaba says audio input costs more than 98% less per hour than the previous generation and audio-video input more than 93% less; see the Bailian platform for current list prices.

Summary

Qwen3.8-Omni-Flash is the natively omnimodal model Qwen released on September 17, 2026, bringing text, image, audio and video input into one model with a 1M token context window. Alibaba's figures put the average score across 29 benchmarks more than 25% above the previous generation, alongside hourly price reductions of more than 98% for audio input and more than 93% for audio-video input. For teams handling long audio and video who want the model to keep going and call tools after it understands the material, it is the most directly relevant tier in the Qwen family.

Category
AI Platforms
Pricing
Freemium
Tags
Omnimodal · Audio-Video · Agent
Website

Related Tools