Image AI-generated by the author with Google Gemini Flash Image
The world moved today.
Something which I never thought would happen until the end of the year - maybe December 2026 - happened today.
An Open Weights Chinese Large Language Model matched Frontier American LLMs.
And in several benchmarks, beat them.
This is a world-shaking event.
A world-changing event.
Let me introduce you to Kimi K3.
And let me tell you why I believe this is the beginning of the end for American Closed-Source LLMs!
Kimi K3 is the first Chinese model to match American Frontier LLMs - at 70% less cost.
US LLMs are expensive to run - because of the 75% profit margin of Nvidia AI Chips.
The US let domestic companies buy Nvidia chips.
Banned China from using them.
China innovated and built its own silicon at much cheaper rates.
Now they serve frontier-class models, of which Kimi K3 is the first, opening them to the entire world to download and use locally.
The next generation of Chinese model iterations will not just match US frontier models - it will beat them, while still being open.
Demand for American LLMs will collapse.
And the US AI bubble will burst.
The main reason, paradoxically: the US not letting China purchase extremely expensive Nvidia infrastructure.
Read on for the full explanation!
On July 16, 2026, Beijing-based Moonshot AI released Kimi K3 - a 2.8-trillion-parameter, open-weight, multimodal reasoning model.
The largest open-source model ever built.
The headline specs are staggering:
Every American AI company should have declared a Code Red, internally and externally, today!
Critically, in a blind test on Web Development, Kimi K3 output was preferred over all other Frontier LLM Models:
Source: https://arena.ai/leaderboard/code/webdev
Now, let me be honest, and separate the hype from the facts.
On aggregate intelligence, K3 does NOT beat the American champions everywhere, but it does reach touching distance.
The Artificial Analysis Intelligence Index v4.1 puts Kimi K3 at 57, behind Claude Fable 5 (60) and GPT-5.6 Sol (59), and just ahead of Claude Opus 4.8 (56).
Fourth configuration, effectively the third-best model family on Earth, on this benchmark (note the caveat - it will be important soon!).
However, it is ahead of Gemini, Grok, and even Opus 4.8, as already mentioned.
And it will soon be available to run locally, with quantization and sufficiently powerful hardware!
Available at: https://the-decoder.com/wp-content/uploads/2026/07/aa_kimik3_index-scaled.jpg
Chart: Artificial Analysis Intelligence Index v4.1 - Kimi K3 scores 57, fourth overall behind Claude Fable 5 (60) and GPT-5.6 Sol (59). Image: Artificial Analysis, via The Decoder.
But look at the individual benchmarks.
Across Moonshot's 35-benchmark launch suite, K3 took first place roughly seven times.
And in real-world task automation, K3 ranked FIRST in four out of eight benchmarks - including AutomationBench, SpreadsheetBench 2, and BrowseComp - beating Claude Fable 5 and GPT-5.6 Sol head-on.
Available at: https://the-decoder.com/wp-content/uploads/2026/07/kimik3_agent_benchmark.webp
Chart: Among general agent benchmarks, Kimi K3 wins three of six tests - including its first-place finishes on automation-style tasks - while Fable 5 leads the visual agent tests. Image: Kimi (Moonshot AI), via The Decoder.
At all times, it stayed within the top two on the benchmarks.
Artificial Analysis's own AutomationBench-AA - their version of Zapier's agentic SaaS workflow evaluation - currently has Kimi K3 leading the entire board at 53 percent.
Not a Chinese lab's self-reported number.
An independent American analytics firm's number!
And in coding, K3 won two of six programming benchmarks outright in the launch suite:
Available at: https://the-decoder.com/wp-content/uploads/2026/07/kimik3_general_benchmark.webp
Kimi K3 wins two of six programming benchmarks and finishes second or third in the rest. Image: Kimi (Moonshot AI), via The Decoder.
There's more:
This, with the model being locally available soon (speculatively), and one unfortunate company spending 500M USD for Claude subscriptions in one month(as per Axios - no cap put on AI token limits by that company)
The main reason for this is the research breakthroughs, which I will cover.
Seven first-place finishes.
Against the four most powerful AI systems on the planet.
From an open-weight model.
Soon running locally with enough unified memory-linked systems.
The world did not just change - it moved from the USA to China.
With the way the tokenomics are working out for all major American Companies, especially OpenAI -
This is the beginning of the end of American domination of the AI space.
Code Red.
Code Red.
Code Red!
Another major release is all it needs for China to cross US expertise and land first in the AGI race -
The US will lose the Manhattan AGI race -
Thanks to the Nvidia ban, China will profitably undercut all US models -
With the open weights release, all major companies will run Kimi K3 locally -
And demand for OpenAI and Anthropic will decrease. Substantially.
Effectively bursting the AI bubble!
This is headline news!
How does a 2.8T model even run economically?
Three genuine innovations:
1. Kimi Delta Attention (KDA). A new attention architecture delivering up to 6.3X faster decoding at million-token contexts, with prefill caching that makes K3's long-context serving commercially viable.
2. Attention Residuals. Boosts training efficiency by roughly 25% while adding under 2% compute overhead. The very technique K3 then out-optimized Fable 5 on in the kernel arena. Poetic, isn't it?
3. Extreme sparsity plus context compaction.
And the economics:
Moonshot reports a 90%+ cache-hit rate on coding workloads.
Artificial Analysis measured $0.94 per task - close to GPT-5.6 Sol ($1.04), half of Opus 4.8 ($1.80), a third of Claude Fable 5 ($2.75)!
Available at: https://the-decoder.com/wp-content/uploads/2026/07/aa_kimik3_index_price-scaled.jpg
Chart: At $0.94 per task, Kimi K3 matches GPT-5.6 Sol's price range at frontier-adjacent intelligence. Image: Artificial Analysis, via The Decoder.
Half of Opus 4.8, beating it on nearly all benchmarks!
Self-reported benchmarks?
Mixed harnesses?
Weights not even released yet?
Here is the honest ledger:
However:
Truth matters.
All of it.
Jeff Geerling's December 2025 testing achieved ~28 tokens/second on Kimi K2 Thinking using a four-Mac Studio M3 Ultra cluster with 1.5TB total memory, drawing under 500 watts versus 5,000+ watts for comparable Nvidia clusters.
Mac Studio clusters are now the first sub-$50,000 setups running trillion-parameter models locally with usable performance. But note the ceiling: those are 512GB Mac Studios.
I strongly expect the demand and the prices to go up steeply in the next few months - and enterprises to be the biggest buyers.
Kimi K3 running locally without quantization for a large company, accounting for the KV cache and multi-user serving, could take 10 Mac Studios with 5 TB unified memory (for enterprises with even less than 100 users, the KV Cache is huge).
Still cheaper than Nvidia and less power and cooling hungry though.
And I see an opportunity for entrepreneurs here to follow Apple, which is already happening.
Systems with computing nodes designed solely for local AI for large enterprises is the upcoming big moat!
Local AI is the future.
The savings are huge.
Especially for OpenClaw and Hermes Agent-like systems and AI agents.
This is definitely the future!
So what does this mean?
First, the open-closed gap has collapsed from 18 months to 0.
Cursor used Kimi to build Composer 2.
DoorDash delegates work to K2.6.
Thinking Machines used K2.5 for post-training data.
American companies are already building on Chinese open models.
Second - the real headline isn't third place.
It's third place at 70% off, and marginal differences from the big boys, and winning first place on several benchmarks, with the promise of a locally running version possible.
No other frontier LLM has a local option, with the exception of GLM 5.2.
For every startup, every solo consultant, every developer in Chennai or Chicago who could never afford $50-per-million output tokens - frontier-class agentic AI just became accessible - with sufficiently powerful hardware (caveat) - that is not Nvidia!
That is a tipping point - THE tipping point.
Third - and this is the biggest one - if China makes Kimi K3 open weights and open source, it will effectively kill the US AI industry.
If the weights ship on July 27 -
If independent hosts in the US and Europe can serve 2.8T economically -
If the agentic wins hold up under same-harness scrutiny -
Then the AI world of August 2026 will look nothing like the AI world of June 2026.
Every gift of intelligence ultimately belongs to all of humanity.
An open model this powerful is not just a threat - it is a victory over the tyranny of Closed Models.
Ai generated by the author with Google Gemini Flash Image.
Sam Altman made the OpenAI investment too big to fail.
With Chinese models running locally, I do not see the demand for OpenAI models anymore.
Anthropic has the constitutional safety moat, which Kimi K3 may or may not have - it’s too early to tell. Being Chinese, I believe it will have.
Google will delay the release of Gemini Pro 3.5 and search for answers.
Grok has open-sourced some parts of its AI model (Grok Build); it would become an incredible win for American AI if it became fully open-source and open-weight.
If Nvidia’s chips did not have a 70% profit margin -
Large Language Models would not be as expensive to run.
What actually forced/helped China to innovate so much and create so many research breakthroughs?
The banning of Nvidia chips in China by the US government.
The irony is not lost on me!
Kimi K3 is a crazy bomb to drop exactly before the IPOs of OpenAI and Anthropic.
However, I do not trust Chinese companies with my proprietary data, so if Kimi K3 cannot be run locally, there is still hope for American AI companies.
Finally (this is rather long term in comparison), within one short year, by July 2027, I don’t see a space for Closed LLMs if China keeps running at this speed.
There was a prediction that Fable 5 class LLMs would run locally by 2028.
Now I believe that mid-2027 is a closer prediction.
I wish I could say, all the best, as I usually do.
But I can’t.
I worry for the USA.
I worry for Anthropic and OpenAI.
I am especially worried about the impact Kimi K3 will have on their IPOs.
However, if China continues to release open-source models:
All the best for the world, and -
OpenAI, Anthropic - I feel your pain.
For the first time since the release of Llama 3, the future of AI is truly open - and:
Although it’s early days to speculate:
Local!
Fable 5-class - and local(with sufficient hardware resources and aggressive quantization)!
AI generated by the author
Around 30% of this article was AI-assisted, but verified and modified substantially by the author.
All images come with their sources linked, or are AI-generated by the author with Google Gemini Flash Image.