Moonshot Introduces 2.8T-Parameter Kimi K3, With Weights Due July 27

Moonshot AI introduced Kimi K3 on July 16 as a 2.8-trillion-parameter mixture-of-experts model with native vision and a 1-million-token context window. The hosted service is live, but Moonshot says the full weights and technical report will arrive by July 27, leaving independent inspection and self-hosted testing pending.
Moonshot AI introduced Kimi K3 on July 16, 2026, describing it as a 2.8-trillion-parameter model for coding, knowledge work and reasoning. The hosted model is available through Kimi's products and API, but Moonshot says the full model weights and a technical report will be released by July 27.
That timing matters. Kimi markets K3 as an open 3-trillion-class model, yet developers cannot fully inspect, modify or self-host the complete model until the promised files arrive.
Architecture described by Moonshot
Moonshot says K3 combines native vision with a 1-million-token context window, Kimi Delta Attention and Attention Residuals. Its sparse mixture-of-experts design activates 16 of 896 experts for each token. The company also says it used quantization-aware training with MXFP4 weights and MXFP8 activations, and recommends supernode configurations with at least 64 accelerators for serving.
The official API documentation lists low, high and max reasoning-effort settings, with max as the default. Moonshot's launch post lists pricing of $0.30 per million cached input tokens, $3 per million uncached input tokens and $15 per million output tokens.
Benchmark results need separation from release facts
Moonshot reports strong results on coding, agentic and knowledge-work evaluations. Those figures use a mix of company-run tests, public leaderboards and differing agent harnesses, so they are not equivalent to a single controlled comparison. Axios reported strong early Arena results but cautioned that hours-old benchmarks and viral demonstrations could overstate real-world reliability. Tom's Hardware likewise noted that the published numbers could not be fully verified before the weights are available.
Nature reported that K3's large size could limit adoption even if its capabilities hold up. That is an infrastructure constraint, not a benchmark result: sparse activation reduces per-token computation, but operators still need to store and serve the full model across a large accelerator fabric.
What practitioners should test next
The July 27 weight release is the main verification gate. Teams should confirm the license, file completeness, memory and communication requirements, quantization quality, long-context degradation and tool-use reliability before treating K3 as a self-hostable production option. Until then, the strongest confirmed facts are the hosted API, Moonshot's disclosed architecture and pricing, and the scheduled release—not reproducible performance of the full downloadable model.
Key Points
- 1Kimi K3 has 2.8 trillion total parameters, native vision and a 1-million-token context window, according to Moonshot.
- 2The hosted service is live, while the full weights and technical report are scheduled for July 27, 2026.
- 3Early benchmark results are promising but remain partly company-reported and cannot substitute for testing the released weights.
Scoring Rationale
Kimi K3 is a technically significant large sparse model with competitive early results and a planned full-weight release. Its practical impact depends on the July 27 artifacts and reproducible independent testing.
Sources
Primary source and supporting public references used for this report.
Practice interview problems based on real data
1,625 SQL & Python problems across 15 industry datasets — the exact type of data you work with.
Try 250 free problems

