Moonshot’s Kimi K3: What You Need to Know About China’s Newest AI Model 🚀
Kimi could be good for the world
My take on the Kimi K3 model that was released last week. To provide some background, Kimi AI models are developed by the Chinese startup Moonshot, which was founded in 2023. The CEO and founder, Yang Zhillin, pursued his PhD at Carnegie Mellon University. He worked for a time at Google Brain and Meta Labs before returning to China. Let’s move on to the model itself. 🤖
I think everyone was genuinely surprised by the model they released last week including me. It is on par with fable 5 and new GPT models. The model is also quite large, containing 2.8 trillion parameters. Where they really excel is cost, as it is significantly cheaper than OpenAI and Anthropic models. Price is important. However, the most unique approach Chinese AI companies usually take is making their models open source. 🔓
Open source (or open weight) usually means you can download the model locally, run it on your system, and even edit the weights. Don’t get me wrong: running a 2.8 trillion parameter model locally is almost impossible. You would essentially need to build a data center to run it. Households cannot really do it, but big companies can, and because it is open weight, they can train it and change the weights. They also don’t have to wonder where their data ends up, since everything can be kept local. 🏢
One thing I noticed is that this model is quite similar to Claude. There is a suspicion that they used a distillation method, basically asking Claude a bunch of questions and training their model accordingly. It was somewhat expected. Chinese companies have really mastered the art of copying. They are so skilled at it that, at some point, a copy can become almost as good as the original. 📋
Why is it generally good to have open source models? I do not believe there should be one entity in the world that we trust to determine what our AI models look like. Also, if only a single market pushes innovation, we face a risk of price control. We usually see that when Anthropic drops a really good model, you see these massive token costs; then, GPT drops something cheaper, and Anthropic needs to adjust. Suddenly, Chinese open source models drop that are almost as good as them, and prices drop even more. I think only China is trying to keep this race going, which is good overall. What I really want to see is more European companies starting to tap into that market. 🌍
Open source models enable you to not have to start from zero; you can build on top of them. This democratizes the technology so we do not have these secret Manhattan Projects. These models are huge, with 2.8 trillion parameters. Would you be able to edit all the weights? Probably not. Will it be biased, with some historical events removed or disregarded? Yes, probably. Will it be favorable to certain groups of people and against others? Yes, probably, but it is still better than starting from zero. 🌱