Kimi K3 Released: Moonshot AI's 2.8 Trillion-Parameter Model, Billed as the First Open 3T-Class LLM

Moonshot AI's Kimi K3 packs 2.8 trillion parameters, an 896-expert MoE design, and a 1M-token context window. Open weights arrive by July 27. Pricing jumps to $3/$15 per million tokens — here's what it means for the AI market.

Kimi K3 Released: Moonshot AI's 2.8 Trillion-Parameter Model, Billed as the First Open 3T-Class LLM

Chinese AI startup Moonshot AI has announced its latest large language model, Kimi K31; the announcement was reported as coming on July 16, 20262. With 2.8 trillion total parameters, the company positions it as “the first open model to reach 2.8 trillion parameters” and “the world’s first open 3T-class model”1. The model weights are scheduled to be released by July 27, 2026 — which would make Kimi K3 one of the largest open-weight models available1.

Kimi K3 features native vision capabilities and a 1-million-token context window1. API pricing is set at $3 per million input tokens (cache miss) and $15 per million output tokens, a substantial increase from Kimi K2.6’s $0.95 input and $4 output pricing2.

Technical Specs: 896-Expert MoE and a New Architecture

Kimi K3 is built on two architectural components developed in-house: Kimi Delta Attention (KDA) and Attention Residuals (AttnRes)1. The company also scaled up Mixture of Experts (MoE) sparsity, effectively activating 16 out of 896 experts1. MoE is a technique that keeps compute costs down by selectively activating only part of a huge model — see our explainer on Mixture of Experts (MoE) for details.

According to Moonshot AI, these structural changes yield an approximate 2.5x improvement in overall scaling efficiency compared to Kimi K21. Further details on the architecture, training, and evaluations will be published in a technical report alongside the weights release1.

Where It Stands: Still Trailing the Top Proprietary Models

Moonshot AI itself states that Kimi K3’s overall performance “still trails the most powerful proprietary models, Claude Fable 5 and GPT-5.6 Sol,” while demonstrating frontier-level performance1.

Third-party results, however, diverge somewhat from that official positioning. In Artificial Analysis’s independent evaluations, Kimi K3 scored an overall Elo of 1547 — up 732 points from K2.6 — trailing only Claude Fable 52. Technology writer Simon Willison also observed that K3 uses 21% fewer output tokens than K2.62.

A Pricing Shift: Among the Priciest Chinese Models

Kimi K3’s API pricing is $0.30 per million tokens for cache-hit input, $3.00 for cache-miss input, and $15.00 for output1. That puts it on par with Anthropic’s Claude Sonnet tier2, and near the top of the range for Chinese labs3. On output pricing, K3 is far cheaper than Claude Fable ($50) but well above z.ai’s GLM-5.2 ($4.40) and DeepSeek V4 ($0.87)3.

Until now, the dominant strategy for Chinese models has been “near-frontier performance at a fraction of the price.” K3’s pricing marks a clear break from that pattern.

2.8x the Scale in One Year

Moonshot AI open-sourced the 1-trillion-parameter MoE model Kimi K2 in July 2025, drawing attention as an agent-oriented model (see our earlier coverage). Exactly one year later, the parameter count has grown 2.8x.

The company’s finances have expanded rapidly as well. In May 2026 it raised $2 billion at a valuation above $20 billion, with annual recurring revenue exceeding $200 million. Its backers include major Chinese tech firms Alibaba, Tencent, and Meituan, along with HSG (formerly Sequoia China), and the company is reportedly preparing for a Hong Kong IPO3. With US-China AI tooling increasingly split — see Alibaba’s restrictions on Claude Code in China — a frontier model backed by Chinese capital carries geopolitical weight.

What It Means: The Ceiling for Open Models Moves Up

For readers evaluating AI adoption, Kimi K3 changes the ceiling of what open-weight models can do. If the weights ship by July 27 as planned, the performance limit for self-hosted models moves close to the frontier. That said, running a 2.8-trillion-parameter model is a serious operational challenge, and most companies will realistically use it via the API.

More options beyond the major US labs also affect pricing leverage. Thinking Machines, led by former OpenAI CTO Mira Murati, recently released its own open model, Inkling, and the gap between frontier and open models keeps narrowing. When selecting a model, it is worth weighing not just performance but data handling, the provider’s capital relationships, and terms of use.

Kimi K3 is available now via the Kimi mobile app, the Kimi API Platform (model name kimi-k3), and Kimi Code in the terminal (selectable via the /model command)1.

Sources

  1. Kimi K3 - Moonshot AI official announcement
  2. Kimi K3, and what we can still learn from the pelican benchmark - Simon Willison’s technical analysis
  3. Moonshot’s Kimi K3 pushes Chinese AI into Fable-level territory - Fortune

We publish the latest AI news every day.

Subscribe via RSS Get new posts the moment they go live.

Search other keywords →