Overview
Kimi is an AI dialogue assistant developed by Moonshot AI, known for its ultra-long text processing capabilities. It was one of the earliest AI products in China to support ultra-long context, attracting attention upon launch. It has since evolved from a simple long-text assistant into a comprehensive AI platform covering Agent, programming, and enterprise office (Kimi Work).
**Major Release: On the evening of July 16, 2026, the new flagship Kimi K3 was officially launched**—with a total of **2.8 trillion parameters** (surpassing DeepSeek V4 Pro's 1.6 trillion), making it **the world's largest open-weight model**. It adopts a MoE architecture (dynamically activating 16 out of 896 experts) + self-developed KDA hybrid linear attention mechanism (Kimi Delta Attention) + attention residual technology, native visual understanding + **1 million token context**. It **topped the Frontend Code Arena globally** (1679 points, surpassing Claude Fable 5), and is now available on Kimi App, Kimi Work, Kimi Code, and API; **complete model weights are planned to be open-sourced on July 27, 2026**. **Notably, the official blog post honestly stated that "overall capabilities still lag behind Claude Fable 5 and GPT-5.6 Sol"** (especially the HLE reasoning shortcoming), a rare display of candor. The day after the release (July 19), reports emerged that Moonshot AI is **expected to complete a Hong Kong IPO within 6 months**.
**Earlier commercialization acceleration**: By mid-June, Kimi's annualized recurring revenue (ARR) exceeded **$300 million** (breaking $100 million in March, $200 million in May). **On June 30, a $20 billion valuation financing round closed, with a new round immediately initiated at a pre-money valuation of $31.5 billion**, raising over $3.9 billion in six months. Technologically, it had previously iterated to **K2.6** (13-hour continuous coding + 300 Agent cluster) and **K2.7 Code** (specialized optimization for programming scenarios).
Kimi excels in scenarios such as paper analysis, long document summarization, and web content extraction. Basic features on the web version and App remain free, while heavy Agent/API scenarios have entered the phase of large-scale commercial monetization.
Key Features
- 200,000-character ultra-long context: A single conversation can handle approximately 200,000 Chinese characters, enabling complete reading and analysis of entire books and lengthy reports
- Kimi K3 (New, World's Largest Open-Source Model): 2.8 trillion parameters, KDA hybrid linear attention + 1 million token context, #1 on Frontend Code Arena globally (1679 points surpassing Fable 5), complete weights open-sourced on July 27
- K2.6 / K2.7 Code Open-Source Models: 13-hour continuous coding + 300 Agent cluster orchestration capability, open-sourced under MIT license; K2.7 Code is specifically optimized for programming scenarios
- Web and File Parsing: Supports direct URL input to extract web content, and uploading PDF/Word/Excel files for in-depth analysis
- Academic Paper Assistant: Proficient in paper reading, literature review, and research methodology analysis, making it a powerful tool for academic researchers
- Online Search: Supports real-time internet search, can retrieve the latest information and provide source links
- Kimi+ Agents: Built-in multiple preset agents (translation, writing, programming, etc.), optimized for specific scenarios
- Data Analysis: Supports uploading data files for analysis and visualization, generating charts and statistical reports
Use Cases
- Academics and business professionals who need to read and analyze lengthy papers, reports, and contracts
- Information workers who frequently need to summarize web content and extract key information
- Students: thesis writing assistance, literature search, knowledge organization
- Editors and content creators who need to process large amounts of Chinese text
- General users looking for a free and powerful Chinese AI assistant
Pros
- Long text processing is a core advantage: ultra-long context, suitable for deep reading and analysis
- Web/App basic features are free: zero cost for daily conversations and document analysis
- Excellent Chinese understanding: accurate and natural grasp of Chinese context
- Convenient file and web parsing: just drop a link or file for analysis
- Solid commercialization validation: ARR grew from $100 million to $300 million in six months, valuation of $31.5 billion, high recognition from capital markets
Pricing
Kimi's web version and mobile app offer basic conversation and file analysis features for free. K3 API pricing: input $3.00/M token (reduced to $0.30/M with cache hit), output $15.00/M token, a significant increase from K2.6 (input $0.95/M, output $4.00/M), corresponding to the 1 million token ultra-long context and stronger reasoning capabilities. Kimi Work (enterprise office scenarios) has a separate commercial plan. With the upgrade of models like K3/K2.6/K2.7, the revenue share from heavy Agent/API scenarios continues to increase.
Summary
If your core need is processing long texts—reading papers, summarizing reports, analyzing contracts—Kimi remains the top choice among domestic AI products, with basic features on the web and app staying free. **The Kimi K3 released on July 16, 2026, is a landmark event in the current round of domestic large model competition**—2.8 trillion parameters making it the world's largest open-source model, top global frontend coding ability, and the official's rare admission that overall capabilities still lag behind Fable 5 / GPT-5.6 Sol, a commendable display of honesty. Kimi is no longer just a "free long-text tool": with $300 million ARR, a $31.5 billion valuation, and a pending Hong Kong IPO, K2.6/K2.7 Code/K3 cover different scenarios. For daily use, it is recommended to pair with Doubao: Kimi for long-text deep tasks and programming agents, and Doubao for daily conversations and light creation.
Version History
- Kimi K3模型已正式上线,定位为专为智能体编程与知识工作打造,现有评测仅提及K3开源商用抽成规则,未覆盖其正式发布及: Kimi K3 model has been officially launched, positioned as specifically designed for agent programming and knowledge work. Existing evaluations only mention K3's open-source commercial royalty rules, without covering its official release and product positioning.
- 官网显示Kimi已上线K3模型,定位为专为智能体编程与知识工作打造,而现有评测仅提及K3开源商用抽成规则,未覆盖K3正式: The official website shows that Kimi has launched the K3 model, positioned as specifically designed for agent programming and knowledge work. However, existing reviews only mention the K3 open-source commercial royalty rules, without covering the official release of K3 and its product positioning, functional features (such as agent programming and knowledge work scenarios), and other new developments.
- K3开源商用抽成规则 (2026-08-07): K3设30%商用收入抽成门槛,Qwen3.8-Max部分基准测试超越K3;全球开发者持续部署K3生态
- 官网主推K3模型(定位智能体编程与知识工作),并新增Kimi Work、Kimi Code、Kimi Claw等产品,评: The official website highlights the K3 model (positioned for agent programming and knowledge work), and has added products such as Kimi Work, Kimi Code, and Kimi Claw, which were not mentioned in the review.
- 官网首页主推K3模型,定位为智能体编程与知识工作,并新增Kimi Work、Kimi Code、Kimi Claw等产品: The official website homepage highlights the K3 model, positioned for agentic programming and knowledge work, and adds entry points for new products such as Kimi Work, Kimi Code, and Kimi Claw, along with plugin and scheduled task features. Existing reviews only mention collaboration between WorkBuddy and Tencent Docs, without covering K3 or the aforementioned new features.
- WorkBuddy 深度打通腾讯文档,打造 Agent 时代的第三代办公协同 (2026-07-30): WorkBuddy's new version deeply integrates with Tencent Docs, supporting direct editing of generated files in the sidebar and embedding Tencent Docs into workflows to enable collaboration between humans, AI, colleagues, and their respective Agents. Users can choose models such as DeepSeek V4 Pro and Kimi K3 to process documents, and upload files to the cloud for sharing and circulation. The author believes that this collaboration model of "humans, Agents, and shared state" marks the arrival of the third generation of office products.
- Kimi K3 开源:2.8T MoE 模型与技术报告 (2026-07-27): Kimi releases its strongest model, Kimi K3, a 2.8T-parameter MoE model with native visual understanding and a 1M token context window. The new architecture achieves a 2.5x intelligence improvement per unit of computation. In addition to model weights, Kimi also open-sources high-performance attention kernels, MoE communication libraries, and large-scale agent runtime environment infrastructure.
- Kimi K3 开源分布式智能体环境 AgentENV (2026-07-27): We have open-sourced AgentENV in collaboration with kvcache-ai. AgentENV is a distributed system for running agent environments at scale. Its components support the reinforcement learning training of the Kimi K3 agent, featuring fast snapshot, recovery, and branching capabilities, making it suitable for large-scale parallel agent workflows. Explore on GitHub: http://github.com/kvcache-ai/AgentEnv
- Kimi K3 开放日:模型权重、技术报告和关键 Infra 技术同步开放 (2026-07-27): Moon's Dark Side releases the 2.8 trillion parameter mixture-of-experts model Kimi K3, supporting native visual understanding and a 1 million token context window. Its scaling efficiency is 2.5 times higher than Kimi K2.5, and it simultaneously open-sources model weights, a technical report, and three Infra technologies: MoonEP, FlashKDA, and AgentEnv.
- Kimi 发布视觉感知基准 PerceptionBench (2026-07-27): Kimi.ai released PerceptionBench, a visual perception benchmark derived from the failure patterns of current frontier models across 42 benchmarks. This benchmark deconstructs visual perception into 10 atomic capabilities and constructs 3,000 validation questions, each testing only a single perceptual ability without requiring reasoning or external knowledge.
- 在 M1 Max 上运行 2.8T 参数的 Kimi K3:Deltafin 项目实现 0.0687 token/s 推 (2026-07-29): The Deltafin project successfully ran the 2.8T-parameter MoE model Kimi K3 on a 64 GB M1 Max, with a current median inference speed of 0.0687 tokens/s (14.6 seconds/token). A full installation requires approximately 1.7 TB of local disk space, while streaming mode needs only 215 GB but reduces inference speed to over 3 minutes/token. The project provides an OpenAI-compatible API server supporting chat and code completion, but it is recommended to set the client timeout to the hour level.
- Kimi K3 Release (World's Largest Open-Source Model) (2026-07-16): 2.8 trillion parameters, world's largest open-weight model, KDA hybrid linear attention + 1 million token context + native visual understanding, #1 on Frontend Code Arena globally (1679 points surpassing Fable 5); official honestly states overall capabilities still lag behind Fable 5/GPT-5.6 Sol; complete weights open-sourced on July 27; July 19 reports Moonshot AI plans Hong Kong IPO within 6 months
- $31.5 Billion Valuation Financing + K2.7 Code (2026-06-30): $20 billion valuation financing closed, new round initiated at pre-money valuation of $31.5 billion; ARR exceeds $300 million; K2.7 Code programming model released, K3 planned for July release
- Kimi K2.6 (2026): Moonshot AI releases K2.6 (MIT open-source), 13-hour continuous coding + 300 Agent cluster; web/app default to K2.6
- Kimi K2 / K2.5 (2025): Transition from long-text assistant to Agent + coding direction
- Early Kimi Chat (2023-2024): One of the first domestic AI dialogue products featuring 2 million character ultra-long context