ChengRang

Qwen3.8-Flash-Next

AI Coding Open Source

Alibaba Tongyi Qianwen, open-sourced late at night on 2026/8/26, is a pioneer preview model based on the next-generation Qwen4 architecture: multimodal MoE, with a total main model parameter of 125B, activating only 6B per token, plus a 51B N-gram vocabulary layer, natively supporting 256K context expandable to 1M; training cost is nearly 90% lower than Qwen3.7-Plus, with weights open-sourced on Hugging Face and ModelScope

AlibabaQwenOpen SourceMoEMultimodal
Visit Qwen3.8-Flash-Next

Disclaimer: Review content represents our editorial team's views and experience, not commercial recommendation or investment advice. Product info and pricing may change; refer to official sources.

Overview

Qwen3.8-Flash-Next is a multimodal MoE model open-sourced by Alibaba Tongyi Qianwen late at night on August 26, 2026, Beijing time. As a preview version of the next-generation Qwen4 architecture, it releases architectural changes ahead of time for global developers to test. The main model has a total of 125B parameters, with only 6B activated per token, and is equipped with an N-gram embedding layer of approximately 51B parameters. It natively supports a 256K context, expandable to 1M, and includes a built-in vision encoder. Its training cost is nearly 90% lower than Qwen3.7-Plus, and the weights have been open-sourced on Hugging Face and ModelScope, with SGLang providing Day-0 support.

Key Features

Use Cases

Pros

Pricing

The sibling model Qwen3.8-Flash (non-Next version) is available on QwenCloud/Qianwen platform as an API service, priced at 1 yuan per million input tokens and 3 yuan per million output tokens (approximately $0.16/$0.47). Qwen3.8-Flash-Next is an open-source weight model, and specific usage costs should be referenced from the official website.

Summary

Qwen3.8-Flash-Next is an open-source precursor model by Alibaba Tongyi Qianwen paving the way for the Qwen4 architecture. With a highly efficient MoE design of 125B total parameters and 6B activated, combined with innovations like hybrid attention and N-gram embedding, it achieves low training costs and local deployment. It performs impressively on coding and inference benchmarks, and through open-sourcing, it promotes community ecosystem adaptation, laying the foundation for the full Qwen4 release.

Version History

Category
AI Coding
Pricing
Open Source
Tags
Alibaba · Qwen · Open Source
Website

Related Tools