Notebookcheck Logo

What's really behind Kimi K3's record numbers

Home screen of the Kimi web interface with input field and K3 logo
ⓘ Screenshot: Kimi.com / Moonshot AI
Behind the simple interface sits the largest open AI model
China’s new flagship model is breaking size records, and the headlines are coming thick and fast. Moonshot itself ranks Kimi K3 behind Claude and GPT. We’ll break down what the numbers really mean, whether you’ll ever be able to run the model yourself, what it actually costs, and what happens to your input.

Kimi K3 has been online since July 16, and the headlines make it sound as if China has just overtaken the US. With 2.8 trillion parameters, it is the largest model whose weights are set to be opened up. Some observers are already calling it the next DeepSeek moment.

The most interesting sentence about this model is not in the headlines, though. It sits in Moonshot's own technical blog. There the maker writes that overall performance still trails the strongest proprietary models, namely Claude Fable 5 and GPT-5.6 Sol. A company that presents its own product and says it is not the best is a rare thing. Which is exactly why the record numbers deserve a closer look.

We have already covered the plain news that K3 is here and free to try. This piece is about what comes after that.

First place, and what it really measures

K3 does win one important comparison. In the Frontend Code Arena, where people see two results side by side and pick the better one without knowing which model produced it, K3 comes out on top. Ahead of Claude and GPT, too.

The catch: this test measures taste. Users rate which web interface they like better. That says a lot about design and little about correctness. Anyone who turns this into "Kimi beats the US models" has missed the difference between a preference ranking and an accuracy test.

Why benchmark tables need careful reading

Moonshot publishes its measurements, and the interesting part sits in the footnotes. Depending on the benchmark, each model ran in a different agent harness, sometimes KimiCode, sometimes Claude Code, sometimes Codex. Those are different starting conditions, not a clean race on the same track.

In one of the coding tests, Moonshot itself concedes that Claude Fable 5 ran into fallbacks in 35 percent of the tasks, which may have dragged its score down. In another test, K3 only reaches its top score with a technique that compacts the context at 300,000 tokens. Without that context management, the number drops.

How easily such figures flip is shown by an evaluation from the independent analysis firm Artificial Analysis. There, Kimi K3 turns up at 49 percent in a hallucination metric. That sounds like a weak score. The scale, however, is inverted: it counts how often the model does not get things wrong. Higher is better, and on the same list Claude Fable 5 sits below K3 at 45 percent, with several GPT-5.6 variants far behind.

A rule worth keeping: before you believe a benchmark headline, check what exactly was measured and which way the scale points.

Can you run K3 yourself?

For our readers, that is the real question. The honest answer: for regular users this will make practically no difference, even once the weights are out. The reason is simply size.

Do the math for a moment. Moonshot trains K3 from the start in a compact number format called MXFP4, which uses about four bits per weight. Weights, also called parameters, are the learned numerical values an AI model is made of. At 2.8 trillion parameters, that puts you at roughly 1.4 terabytes for the model file alone. Not gigabytes. Terabytes. For perspective: a typical model with eight billion parameters, the kind that runs on a well-equipped laptop, takes up around five gigabytes.

A 1.4-terabyte model file: K3 is far too big for any home PC

Moonshot itself recommends running K3 on supernode setups with 64 or more accelerators. That is not excessive caution but the plain consequence of these dimensions. A single professional graphics card with 96 gigabytes of memory is off by more than a factor of ten.

"Open weights" therefore does not mean you will run this model at home. It means researchers, companies and cloud providers may download it, inspect it and run it on their own infrastructure instead of only renting it through Moonshot's servers. That is valuable, but it is something different from what many people picture.

If you actually want AI without the cloud on your own machine, the small models built for that job are the right choice. We have shown how that works and what your hardware needs in a separate article.

A side note for the technically curious: because K3 was trained at this low precision from day one, the compact format is the original state, not a model squeezed down after the fact. The usual question of how much quality the shrinking cost does not arise in the same way.

What it costs in practice

Billing works in tokens, small text chunks of usually just a few characters. Through the API, Moonshot charges $3 per million input tokens and $15 per million output tokens. Repeated inputs are billed at 30 cents, which is far cheaper.

That makes K3 no bargain anymore. It costs about three times as much as its own predecessor, which charged just under a dollar for input. Against Claude Fable 5 at $10 and $50, K3 is clearly cheaper. Claude Sonnet 5, though, still undercuts it at its current introductory price of $2 and $10 and only matches K3 in September. The era of Chinese models competing mainly on price is ending at Moonshot.

K3 costs three times as much as its predecessor but stays below Fable 5

There is one more point that is easy to miss. K3 currently always thinks at its highest effort level. Moonshot plans to add leaner modes later. That makes answers thorough, but it also means even a simple question triggers the expensive reasoning work. Independent measurements also clock the model at around 62 output tokens per second, mid-pack for speed.

What happens to your inputs

For anyone trying K3 in the app or on the website, this is the section that matters most. Moonshot's privacy policy states plainly that inputs, audio, images, videos and files are processed to provide and improve the services, explicitly including the training and optimization of its own models.

A dedicated switch to turn off training on your content is nowhere in the policy. It points to general data subject rights, which you can exercise through your account settings or by emailing the privacy contact. ChatGPT and Claude, by contrast, put a toggle right in the settings.

Also in the policy: collected data includes IP address, device identifiers, session and conversation IDs and, where device settings allow it, data from the clipboard. On storage location, the policy only says data may be transferred to servers outside your country of residence. No country is named. The provider is Moonshot AI PTE. LTD., based in Singapore, and the policy is dated July 7, 2025.

For a test run with harmless tasks, that is acceptable. For company internals, customer data or private documents, the same rule applies as with any cloud service, only here with a provider whose own terms explicitly provide for training.

Who gets the most out of Kimi K3

Trying it costs nothing. Two groups stand to gain the most: anyone who develops or designs, and anyone who regularly works with very long documents. K3 is strong where code and visuals meet, in web interfaces, 3D and animation. The one-million-token context window is a real argument for long texts, and image processing is built in.

If you want the best answer quality or the smoothest experience, the top US models remain the better choice. Moonshot says so itself. And anyone hoping to put a frontier model on their own computer soon should keep those 1.4 terabytes in mind.

The remarkable thing about K3 is less the ranking than the pace. A Chinese company is delivering a model on a par with the world's best and is opening up the model data. Two years ago, that would have been hard to imagine.

Google LogoAdd as a preferred source on Google
Mail Logo

No comments for this article

Got questions or something to add to our article? Even without registering you can post in the comments!
No comments for this article / reply

static version load dynamic
Loading Comments
Comment on this article
> Expert Reviews and News on Laptops, Smartphones and Tech Innovations > Reviews > What's really behind Kimi K3's record numbers
Steffen Zahn, 2026-07-23 (Update: 2026-07-19)