Overview
Ant Group Bailing Ling-3.0-flash is a new generation native hybrid reasoning model released by Ant Group Bailing, with a total parameter count of 124B and an activated parameter count of only 5.1B. It matches or even surpasses the previous flagship Ring-2.6-1T in metrics such as traditional reasoning, instruction following, and long text processing. The model adopts a native hybrid linear attention architecture and 1/64 sparse MoE, reducing TTFT by 60% to over 80% under long inputs.
Key Features
- Native Hybrid Reasoning: Total parameters 124B, activated only 5.1B, matching the previous flagship Ring-2.6-1T
- Sparse MoE Architecture: 1/64 sparse MoE plus native hybrid linear attention
- Low Latency for Long Text: TTFT reduced by 60% to over 80% under long inputs
- Multi-Environment Training: Expanded to 10,000+ interactive training environments
Use Cases
- Enterprise applications requiring efficient reasoning
- Long text processing scenarios
- Real-time applications sensitive to latency
Pros
- Extremely low activated parameters, high reasoning efficiency
- Significantly reduced latency for long text
- Performance matching the previous flagship
Pricing
API pricing is subject to the official Ant Group Bailing website.
Summary
Ling-3.0-flash matches the previous 1T flagship with only 5.1B activated parameters, representing a masterpiece of sparse MoE and hybrid reasoning, suitable for enterprise scenarios with requirements for reasoning efficiency and long text latency.
Version History
- Ling-3.0-flash released (2026-07-24): 124B parameters, activated 5.1B, native hybrid linear attention plus 1/64 sparse MoE, TTFT reduced by 60-80% under long inputs