Overview
Ant Bailing Ling-3.1-flash is a MoE large model released by Ant Group's Bailing large model (InclusionAI) on September 30, 2026. The model has about 560B total parameters, activates about 25B parameters per token, and has a context window of up to 1M, allowing it to accommodate longer documents, code, and task histories. The official description states that during development it has been continuously optimized around tasks such as general agents, search, daily office work, and software development, while also continuously improving capabilities in healthcare, finance, and materials science research scenarios. Its previous generation, Ling-3.0-flash, had 124B total parameters and 5.1B activated parameters, was open-sourced under MIT, and used a MoE architecture with a 5:1 hybrid of Kimi Delta Attention and Mamba-like linear attention, with 512 routed experts. Ling-3.1-flash offers a two-week free trial, with a service length of 256K during the trial period, after which it becomes paid and is planned to be open-sourced.
Key Features
- Large-scale MoE architecture: About 560B total parameters, with about 25B parameters activated per token, controlling the amount activated per inference while maintaining a relatively large parameter scale.
- 1M context window: The context window upper limit is 1M (1 million tokens), allowing it to accommodate longer documents, code, and task histories.
- Optimization for multiple scenarios: The official description states that it has been continuously optimized around tasks such as general agents, search, daily office work, and software development, while also continuously improving capabilities in healthcare, finance, and materials science research scenarios.
- Free trial and subsequent opening plans: A two-week free trial is offered, with a service length of 256K during the trial period; after the free trial period ends and it transitions to a paid service, it is planned to open the 1M context, and it is also planned to be open-sourced at the same time and continue updating model performance.
- Continuation of the previous generation's technical route: The previous generation Ling-3.0-flash used a MoE architecture with a 5:1 hybrid of Kimi Delta Attention and Mamba-like linear attention, with 512 routed experts, providing a foundation for the development of Ling-3.1-flash.
Use Cases
- General agent tasks: The official description states that it has been continuously optimized around the general agent direction and can support agent scenarios requiring multi-step planning and tool calling.
- Search scenarios: The official description states that it has been continuously optimized for search tasks, making it suitable for applications requiring retrieval and information integration.
- Daily office work: The official description states that it is optimized for daily office tasks and can assist with document processing, information organization, and similar work.
- Software development: The official description states that it has been continuously optimized in the software development direction and can assist with code understanding, generation, and task history management.
- Professional domain research: The official description states that it continuously improves capabilities in healthcare, finance, and materials science research scenarios and can serve analysis and research tasks in related fields.
Pros
- Large parameter scale: About 560B total parameters, the largest parameter scale among domestic Flash-series MoE models in the same period.
- High context window upper limit: Supports a 1M context, allowing it to accommodate longer documents, code, and task histories.
- Reasonable control of activated parameters: About 25B parameters activated per token, maintaining a relatively controllable activation amount despite the large total parameter count.
- Clear multi-scenario optimization: The official description states that it has been continuously optimized around tasks such as general agents, search, daily office work, and software development, and also covers professional scenarios such as healthcare, finance, and materials science.
- Free trial provided: A two-week free trial is offered, with a service length of 256K during the trial period, making it convenient for users to try it first.
- Clear subsequent opening plans: After the free trial period ends and it transitions to a paid service, it is planned to open the 1M context, and it is also planned to be open-sourced at the same time and continue updating model performance.
Pricing
Ling-3.1-flash offers a two-week free trial, and during the free trial period the model service length is 256K. After the free trial period ends and it transitions to a paid service, it is planned to open the 1M context, and it is also planned to be open-sourced at the same time and continue updating model performance. As of now, the official side has not announced specific pricing figures, and the pricing section is subject to the official website.
Summary
Ant Bailing Ling-3.1-flash is a MoE large model released by Ant Group's Bailing large model on September 30, 2026, with about 560B total parameters, about 25B activated per token, and a context window upper limit of 1M. The official description states that it has been continuously optimized around tasks such as general agents, search, daily office work, and software development, and has improved capabilities in scenarios such as healthcare, finance, and materials science. The model offers a two-week free trial, with a service length of 256K during the trial period, after which it becomes paid and is planned to be open-sourced. As of now, the official side has not published public benchmark comparisons or specific pricing, and there is no reliable basis for horizontal ranking for the time being; pricing is subject to the official website.
Version History
- Ling-3.1-flash release (2026-09-30): Ant Group InclusionAI released Ling-3.1-flash, a MoE model with about 560B total parameters and roughly 25B activated per token and a 1M context window, tuned for general-purpose agents, search, everyday office work and software development, with gains in medical, finance and materials science scenarios. It opens with a two-week free trial capped at 256K service length, after which the 1M context becomes available under paid service and open-sourcing is planned