Tekshove

← View all articles
AI

Kimi K3: The Ultimate Guide to China’s 2.8T AI Model

A #1 ranking on one chart. A 51% number missing from another. Kimi k3, in a single image.

Kimi k3 launched on July 16, 2026, and immediately grabbed attention. A 2.8-trillion-parameter open-weight model from China’s Moonshot AI, it beat Claude Fable 5 in blind developer testing and landed near the very top of the coding leaderboards. Then, a second story emerged, one Moonshot’s own charts never mentioned. In short, this guide covers what kimi k3 actually does well, what remains unverified, and the specific gap between the launch pitch and independent testing.

The short answer
Kimi k3 is a 2.8-trillion-parameter open-weight model from Moonshot AI, released July 16, 2026. It scores close to the frontier on coding benchmarks and topped a blind developer test ahead of Claude Fable 5. However, independent testing found a 51% hallucination rate that Moonshot’s own launch charts left out. New subscriptions were paused on July 20 due to a capacity crunch, and full open weights are not expected until July 27, so every number remains unverified until then.

1.What is Kimi K3?

Kimi k3 comes from Moonshot AI, a Beijing-based lab. Moonshot calls it the largest open-weight model released so far, at 2.8 trillion parameters. It uses a mixture-of-experts design with 896 expert parts, turning on just 16 per token. That design keeps running costs low despite the model’s huge total size. Moonshot built it for coding, knowledge work, and long-running agent tasks. Its context window stretches up to one million tokens.

2.The benchmark claims

On paper, the numbers are genuinely strong. Kimi k3 scored 88.3% on Terminal-Bench 2.1. That is just behind GPT-5.6 Sol’s 88.8%, and ahead of both Claude Opus 4.8 and GPT-5.5. On Frontend Code Arena, meanwhile, blind developer votes ranked it first overall, ahead of Claude Fable 5. Moonshot itself is more measured than the headlines. The company says K3 still sits behind Fable 5 and GPT-5.6 Sol on overall performance. Still, it beats every other major model in its own evaluation suite.

kimi k3 in 7 factsKimi k3, separated into seven confirmed and disputed facts.

3.The number missing from the launch charts

Here is the part that matters most. Independent testing found kimi k3 has a 51% hallucination rate. Moonshot’s own launch benchmarks never showed this number. A model that tops a coding leaderboard while getting facts wrong half the time is not really a contradiction. Coding tests and general fact-checking measure two very different things. Still, leaving out a bad number while sharing every good one is a real honesty problem. It is not a small detail to skip.

Trying to figure out which AI model is actually right for your team?

We track model releases against real production use, not launch-day marketing. Instead, let’s talk through what genuinely fits your workflow.

Talk to our team →

4.The capacity crunch

Demand outran Moonshot’s infrastructure almost right away. New Kimi subscriptions were paused on July 20, just four days after launch, because demand strained available capacity. Running kimi k3 yourself is not realistic for most businesses either. Moonshot recommends 64 or more accelerators in a single high-bandwidth cluster, plus roughly 1.4 terabytes of fast memory. That needs NVIDIA Blackwell-class or AMD MI400-class hardware. In short, no consumer or small-business GPU setup clears that bar.

5.The distillation dispute

This launch also reopened an older argument. Back in February 2026, Anthropic accused Moonshot of using 3.4 million Claude exchanges to train its models through distillation. Kimi k3 now scores within a few points of the exact models named in that dispute. Notably, Moonshot has not addressed the accusation directly in its K3 launch materials. This context matters for anyone judging the model’s true origins, apart from how well it performs.

6.Pricing and access

Kimi k3 is available now through Kimi.com, the official API, OpenRouter, and Cloudflare Workers AI, ahead of its planned open-weight release. Pricing runs $0.30 per million cache-hit input tokens, $3 per million on cache misses, and $15 per million output tokens. Notably, that output price is more than three times higher than Moonshot’s own K2.6 model.

For everyday production workloads, K2.6 remains the cheaper default. K3 makes more sense reserved for tasks that genuinely need frontier-level coding or massive context. This Tom’s Hardware coverage of the launch and this Tech Times report on the hallucination rate cover the full details.

Model-agnosticWe evaluate every new model against your real workload, not the launch hype
Cost-awareWe help you match model tier to task, since the priciest option is rarely the right default
Verified firstWe wait for independent benchmarks before recommending any new release

What this means for your business

Kimi k3 is a genuine technical achievement, built under real hardware constraints. It is also not yet a verified, safe default for production work. Treat every number in this story as one of two things: independently confirmed, or a claim still waiting on open weights and outside testing. For coding tasks, that distinction matters more than the leaderboard rank. Our guides to the Gemini 3.5 Pro delayClaude Sonnet 5, and Grok 4.5 cover the broader AI model landscape this fits into.


Frequently asked questions

What is Kimi K3?

Kimi K3 is a 2.8-trillion-parameter open-weight AI model from Beijing-based Moonshot AI, released July 16, 2026. It is built for coding, knowledge work, and long-horizon agent tasks, and Moonshot describes it as the largest open-weight model released to date.

Is Kimi K3 actually better than Claude or GPT at coding?

The results are close on several benchmarks. Kimi K3 scored 88.3% on Terminal-Bench 2.1, just behind GPT-5.6 Sol’s 88.8%. It also topped a blind developer evaluation on Frontend Code Arena. However, all of these numbers come from Moonshot’s own reporting or hosted API testing, not yet from independently verified open weights.

Can I run Kimi K3 on my own hardware?

Not on typical consumer or small-business hardware. Moonshot recommends 64 or more accelerators in a single high-bandwidth cluster and roughly 1.4 terabytes of fast memory. This requires NVIDIA Blackwell-class or AMD MI400-class GPUs.

Want help choosing the right AI model for your actual work?

TekShove can help you cut through launch-day hype and pick tools that are genuinely verified and production-ready.

Talk to Our Team

Tell us what you need help with and our team will get back to you.

Need help with your next digital move?

From design to development, we help brands create better digital experiences with clarity, consistency and measurable results.

Let’s connect
Scroll to Top