Moonshot Unveils 2.8-Trillion-Parameter Kimi K3 Model

Moonshot AI unveiled Kimi K3 on July 17, a 2.8-trillion-parameter open-weight model with native visual understanding and a 1 million-token context window. Moonshot's documentation identifies Kimi Delta Attention, Attention Residuals, and a sparse mixture-of-experts design as core architectural components; it lists full model-weight release by July 27. Reuters and CNBC report that the company claims competitive results against leading U.S. systems on selected evaluations.
Moonshot AI unveiled Kimi K3 on July 17, presenting a 2.8-trillion-parameter model aimed at reasoning, long-horizon coding, and knowledge-work tasks. Moonshot's Kimi documentation describes K3 as an open-source model in the 3-trillion-parameter class, with native visual understanding and a 1 million-token context window. The company lists July 27, 2026, as the date by which full model weights are scheduled for release.
The release is notable both for claimed model scale and for the pending availability of weights. Reuters reported that Moonshot called K3 the world's largest open-weight AI system, while noting that open-weight models can be downloaded, run, and customized by users. The distinction matters: an announced model and a fully released model are operationally different milestones for teams seeking to evaluate weights and inference behavior.
Architecture and claimed capabilities
Moonshot's documentation attributes K3's architecture to Kimi Delta Attention (KDA), described as a hybrid linear-attention mechanism, and Attention Residuals (AttnRes). According to the documentation, the system uses the Stable LatentMoE framework and activates 16 of 896 experts, a sparse mixture-of-experts configuration intended to limit active computation relative to the model's total parameter count.
Moonshot claims these architectural and training changes provide roughly 2.5 times K2's scaling efficiency. The company also describes K3 as capable of sustaining long-running engineering work with limited human supervision, operating on large codebases, coordinating terminal tools, and combining software tasks with visual inputs such as screenshots. Those are vendor claims, and the company stated that further architecture, training, and evaluation details would accompany a technical report.
For ML infrastructure teams, the 2.8-trillion total parameter figure should not be read as a direct measure of per-request compute. In sparse MoE systems, the number of activated experts, expert routing, sequence length, precision format, and serving stack are more immediately relevant to latency and memory planning. A 1 million-token context window also creates separate questions around usable long-context quality, attention-memory behavior, retrieval design, and cost under production workloads.
Benchmark claims and external signals
Reuters reported that Moonshot said K3 performed competitively with Anthropic's Fable 5 with fallback on GPU-kernel optimization and outperformed Anthropic Opus 4.8, OpenAI GPT 5.6 Sol, and GPT 5.5 in that evaluation. CNBC separately reported that Moonshot said K3 trailed Fable 5 and GPT 5.6 Sol on overall performance while outperforming other tested systems on benchmarks that included coding and general-agent tasks.
Reuters also reported that Arena.ai ranked K3 first in a web interface-building benchmark and that Vals AI placed it second overall behind Fable 5. Such rankings can be useful directional evidence, but practitioners typically need task-level data, test-set provenance, prompting details, tool configurations, and reproducible serving conditions before translating benchmark positions into production-model selection.
Bank of America analysts, quoted by CNBC, said that K3 demonstrated how pre-training scale and architectural innovation can deliver substantial gains despite China's hardware and compute constraints. More broadly, comparable open-weight releases expand the set of models that organizations can inspect and customize, while shifting evaluation work toward practical concerns such as weight access, hardware compatibility, inference throughput, safety controls, and benchmark reproducibility.
Moonshot's documentation states that it is working with inference partners and open-source maintainers to align technical details ahead of the weight release. Until the weights and technical report are available, independent assessment of K3's reproducibility and real-world serving characteristics remains limited.
Key Points
- 1Moonshot unveiled a 2.8-trillion-parameter sparse MoE model, but full weights are scheduled for release by July 27.
- 2K3 combines native vision and a 1 million-token context window, making long-context quality and serving cost central evaluation questions.
- 3Comparable open-weight releases broaden customization options, while requiring teams to validate weight access, throughput, safety controls, and benchmark reproducibility.
Scoring Rationale
Kimi K3's reported scale, sparse MoE architecture, and planned weight release make it a major model-development event for practitioners evaluating open-weight frontier systems. Its practical impact depends on the July 27 weight release, technical documentation, and independent reproducibility evidence.
Sources
Primary source and supporting public references used for this report.
View 4 more sources
- China's Moonshot unveils world's largest open AI model, closing in on US rivalsreuters.com
- China's Moonshot AI claims Kimi K3 can rival OpenAI and Anthropicbbc.com
- Chinese AI has leveled up, and brought renewed focus on the open weight model shiftcnbc.com
- China delivers a one-two punch to America’s AI dominancetheverge.com
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems
