Moonshot AI introduces the 2.8-trillion-parameter Kimi K3 open-weight model

Moonshot AI introduced Kimi K3, a 2.8-trillion-parameter multimodal model with a 1-million-token context window. The company's technical blog and developer documentation describe a sparse mixture-of-experts architecture, Kimi Delta Attention and support for long-horizon coding and visual tasks. Moonshot says full model weights are scheduled for release by July 27; until then, independent teams cannot verify the complete licensing, checkpoint and deployment profile. Reuters reports strong results in selected third-party evaluations, but those scores do not establish universal superiority. For practitioners, the key questions are reproducible benchmarks, serving memory, throughput, quantization and commercial license terms.
What Moonshot released
Moonshot AI introduced Kimi K3, a 2.8-trillion-parameter multimodal model with a 1-million-token context window. The company currently offers access through Kimi's online and developer services and says full model weights are scheduled for release by July 27, 2026.
That timing matters. Hosted access lets developers test an API, but independent deployment and modification depend on the actual checkpoint, license, tokenizer, model code and supported inference stack. Until those artifacts arrive, K3's full open-weight deployment profile cannot be independently verified.
Architecture described by Moonshot
Moonshot's technical blog and developer documentation describe three notable components:
- •Kimi Delta Attention, intended to improve long-context efficiency.
- •Attention Residuals, used in the model's residual pathway.
- •Stable LatentMoE, a sparse mixture-of-experts design that activates 16 of 896 experts for a request.
Moonshot claims roughly 2.5x scaling efficiency compared with Kimi K2 and positions K3 for long-horizon coding, knowledge work, reasoning and visual software-engineering tasks. Those are vendor-reported claims. A promised technical report with fuller training and evaluation details was not available in the retrieved sources.
| Evidence available now | Evidence still needed |
|---|---|
| Official technical overview and API documentation | Downloadable weights and final license |
| Reported 2.8T scale and 1M-token context | Independent memory, latency and throughput measurements |
| Selected external benchmark placements reported by Reuters | Reproducible evaluation across broader workloads |
| Hosted developer access | Stable self-hosting and quantization guidance |
What the benchmark reports mean
Reuters reports that K3 placed strongly in selected evaluations from Arena.ai and Vals AI and that Moonshot published competitive task-specific comparisons. These results are useful signals, but they should not be treated as a universal ranking. Model performance can change substantially across coding repositories, retrieval tasks, languages, visual inputs, context lengths and tool-use environments.
The New York Times separately reports the launch in the context of increasing competition between Chinese and US AI labs. That industry context does not replace workload-specific validation.
LDS assessment
For practitioners, K3 raises the scale ceiling for the open-weight ecosystem but does not eliminate frontier-scale deployment costs. Sparse activation can reduce per-request compute while the full checkpoint still creates substantial storage, loading, partitioning and serving demands.
Teams should wait for the weights and license, then measure end-to-end latency, concurrency, context-cache behavior, quantization loss and total infrastructure cost on their own workloads. A smaller model that fits established infrastructure may remain the better production choice even when a larger model leads selected benchmarks.
The source record is now curated around Moonshot's official blog and documentation plus two independent reports. Earlier derivative and duplicate references are not presented as separate evidence.
Key Points
- 1Moonshot introduced a 2.8-trillion-parameter multimodal model with a 1-million-token context window.
- 2Full weights are scheduled for July 27, so license and self-hosting claims remain incomplete until release.
- 3Sparse activation may reduce request compute, but storage, memory and serving costs remain material at this scale.
Scoring Rationale
Kimi K3 is a major claimed advance in open-weight model scale and long-context capability. Its practical impact depends on the promised weight release, licensing, independent evaluation and deployability evidence.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems
