Overview
The new generation native hybrid reasoning model Ling-3.0-flash released by Ant Bailing, with total parameters of 124B and activated parameters of 5.1B, achieves extreme intelligence density to benchmark and surpass 1T-level flagship reasoning models. Through underlying architecture upgrades and Agent-oriented optimization, the model efficiently transforms parameters into actual capabilities, demonstrating the potential for cross-level challenges at an extremely low scale.
Key Features
- Hybrid Linear Architecture, Native Evolution: From the beginning of pre-training, it adopts a native hybrid linear attention architecture (5:1 alternating stacking of KDA and MLA), and upgrades to a new KDA fine-grained diagonal gating and 1/64 sparse MoE. With total parameters of 124B and activated parameters of 5.1B, it achieves a synergistic leap in long-context efficiency, state memory, and computational cost.
- Extreme Intelligence Density, Cross-Level Challenge: Pursuing the limits of intelligence density and intelligence efficiency that are fast, economical, and deployable, it achieves performance cross-level challenges when facing larger competitors' SOTA and previous-generation flagships. With only 5.1B activated parameters, it provides excellent traditional reasoning, instruction following, and long-text capabilities, fully empowering efficient deployment of complex Agent scenarios.
- Comprehensive Agent Evolution, Long-Range Closed Loop: Deeply polished around real productivity scenarios, expanding over 10,000 interactive training environments to achieve full closed-loop completion of Coding, General, and Deep Research Agent tasks. At the same time, it innovatively integrates SGLang HiCache + Mooncake hierarchical caching architecture, enabling long-range interactions without recomputation, reducing TTFT by 60% to over 80% for long inputs.
Use Cases
- Complex planning and multi-turn tool invocation
- Long-range task delivery and Deep Research
- Code generation and general Agent scenarios
Pros
- Benchmarks against 1T-level flagship models with only 12.4% total parameters and 8.1% activated parameters, offering extremely high cost-effectiveness
- Native hybrid linear attention and sparse MoE improve parameter conversion efficiency
- Agent optimization covers over 10,000 training environments, ensuring stable closed-loop for long-range tasks
- Hierarchical caching architecture significantly reduces latency for long inputs
Pricing
Specific pricing not disclosed, but emphasizes a balance of high performance and high cost-effectiveness, suitable for enterprise-level Agent deployment
Summary
Ling-3.0-flash is a native hybrid reasoning model launched by Ant Bailing, achieving a breakthrough in intelligence density with 124B total parameters and 5.1B activated parameters. It benchmarks or even surpasses 1T-level models in reasoning, instruction following, and long-text capabilities. Its hybrid linear architecture and Agent-oriented optimization make it excel in complex planning, tool invocation, and long-range tasks, while reducing latency through hierarchical caching, making it a productivity-level model that balances performance and cost.