Overview
Kimi K3 is the new flagship large model officially released by Moonshot AI on the evening of July 16, 2026, and is the company's most capable model to date. K3 has a total of **2.8 trillion parameters** (MoE architecture, dynamically activating 16 out of 896 experts), surpassing DeepSeek V4 Pro's 1.6 trillion, making it the **world's largest open-weight model**. Technically, it adopts the self-developed **KDA hybrid linear attention mechanism** (Kimi Delta Attention) + Attention Residuals technology, natively supports visual understanding, and has a **1 million token context window**.
**Performance highlights**: Topped the Frontend Code Arena with **1679 points**, surpassing Claude Fable 5. It has been simultaneously launched on Kimi App, Kimi Work, Kimi Code, and API, with **full model weights open-sourced on 2026/7/27**.
**Rare honesty**: The official blog homepage explicitly states that "overall capabilities still lag behind Claude Fable 5 and GPT-5.6 Sol," especially in the HLE (Humanity's Last Exam) reasoning test, without only showcasing winning benchmarks like most vendors. This attitude of "daring to display the loss table" has gained considerable recognition in the industry.
The day after release (7/19), reports emerged that Moonshot AI could complete a Hong Kong IPO **within 6 months**, with ARR already reaching $300 million, and K3's release is expected to drive several-fold growth.
Key Features
- 2.8 trillion parameters (largest open-source globally): MoE architecture, dynamically activating 16 out of 896 experts, surpassing DeepSeek V4 Pro (1.6T) to become the world's largest open-weight model
- KDA hybrid linear attention mechanism: Self-developed Kimi Delta Attention + Attention Residuals technology, the core technical support for K3's inference efficiency and long-context capability
- 1 million token context + native visual understanding: Supports ultra-long document/codebase processing, with native image input understanding without additional adaptation
- Frontend Code Arena world first: Topped the frontend code arena with 1679 points, surpassing Claude Fable 5, currently the most recognized strength
- Official proactive disclosure of weaknesses (honest stance): Blog homepage states overall capabilities still lag behind Fable 5 / GPT-5.6 Sol, with significant gaps in HLE reasoning tests, no selective display
- Simultaneous launch on four platforms: Kimi App, Kimi Work, Kimi Code, and API support simultaneously, full weights open-sourced on 2026/7/27
Use Cases
- Frontend code generation and UI development (current strongest area)
- Long document/large codebase analysis requiring million-token context
- Complex tasks combining multimodal (image understanding) and text
- Enterprises and research institutions seeking the world's largest open-source model for self-deployment
- Technical evaluation scenarios comparing domestic open-source models with Fable 5/GPT-5.6
Pros
- World's largest open-weight model, a landmark breakthrough in technical influence
- Frontend Code Arena world first, top-tier frontend code generation capability
- 1 million token native context + visual understanding, solid multimodal capability
- Official proactive disclosure of weaknesses, rare honest stance in the industry, trustworthy
- Simultaneous launch on four platforms (App/Work/Code/API), fast deployment
Pricing
API pricing: Input $3.00/M token (reduced to $0.30/M with cache hit), output $15.00/M token, significantly higher than K2.6 (input $0.95/M, output $4.00/M), corresponding to the 1 million token ultra-long context and stronger reasoning capability. Available directly in Kimi App/Work/Code (free and paid quotas depend on specific product lines). Full open-source weights will be available for free download and self-deployment after 2026/7/27, but due to the massive total parameter size, local personal deployment is essentially impractical; feasible paths are community quantized versions + cloud-hosted inference.
Summary
Kimi K3 is one of the landmark events in the 2026 domestic large model competition—2.8 trillion parameters making it the world's largest open-source model, top frontend code capability globally, and the rare official admission on the release homepage that overall capabilities still lag behind Claude Fable 5 and GPT-5.6 Sol. This honest stance is more commendable than simply stacking parameters. If your core need is frontend code generation or a super-large open-source model for self-controlled deployment, K3 is currently the most noteworthy option; but if you need top-tier general reasoning capability (HLE, etc.), you should still prioritize Fable 5 / GPT-5.6 Sol. After the full weights are open-sourced on 7/27, it is recommended to follow community quantization solutions to lower the deployment threshold.
Version History
- Kimi K3 official release (2026/07/16): 2.8 trillion parameters, world's largest open-weight model; KDA hybrid linear attention + 1M token context + native visual understanding; Frontend Code Arena world first (1679 points, surpassing Fable 5); official proactive disclosure of HLE and other reasoning tests still lagging behind Fable 5/GPT-5.6 Sol; simultaneous launch on four platforms; full weights open-sourced on 7/27
- Kimi K2.7 Code (2026/06/12): Minor iteration focused on coding; 1.1T MoE/32B activated/256K context; surpassed K2.6 on 3 coding benchmarks; token consumption reduced by 30%
- Kimi K2.6 (2026/04/20): Open-sourced under MIT license; 13-hour continuous coding + 300 Agent cluster; coding capability comparable to top international closed-source models at the time