Bullish

Kimi K3 Efficiency Surges 2.5x with 2.8T Parameters and Rewritten Architecture

2026-07-28 00:20:56

Moonshot AI releases Kimi K3 featuring 2.8T parameters and 2.5x scaling efficiency. The model matches Fable 5 in benchmarks via architectural upgrades and specialized post-training.

Woofun AI reports that Moonshot AI has published the technical report for Kimi K3, a model with 2.8 trillion parameters where each token activates 104 billion. The new architecture achieves 2.5 times higher scaling efficiency than K2, requiring only 40% of the computational volume for equivalent validation loss.

The model introduces KDA for long sequences, Attention Residuals to prevent information dilution, and a redesigned MoE with 896 routing experts activating 16 per token. Post-training involves merging nine specialized experts trained on general, Agent, and code tasks. K3 performance now approaches or surpasses Fable 5 and GPT-5.6 Sol in code, search, and tool call benchmarks.

WOOFUN AI

Impact Assessment · Quick Read

Kimi K3’s architectural innovations suggest a shift from pure parameter scaling to efficiency-driven model design. By matching top-tier competitors like Fable 5 with significantly lower compute requirements, Moonshot AI may gain a cost advantage in deploying large-scale Agent capabilities. This development could intensify competition in the high-performance LLM sector, particularly for applications requiring long-context retention and complex tool usage.
Generated by WOOFUN AI · For reference only, not investment advice

Comments

Me
Replying to @User
0/800

No comments yet.

Notifications

Sign in to view messages
View all messagesManage subscriptions