Key Takeaways

- Kimi K3 scored 32.2% on ExploitBench versus 76.2% for leading US frontier models
- The model completed only 17 of 32 steps in a simulated network attack, compared to 28.5 for US models
- Results support US allegations that Moonshot AI distilled capabilities from more advanced Western models
Moonshot AI's Kimi K3 scored 32.2 percent on a benchmark measuring exploit development skills, less than half the 76.2 percent average logged by leading US frontier models. The joint evaluation by the British AI Security Institute and the US Center for AI Standards and Innovation tested the Chinese model's offensive cyber capabilities and found it willing to assist with attacks but unable to execute the hardest exploits.
The results carry weight beyond the raw numbers. They align with US government allegations that Moonshot AI built Kimi K3 by distilling output from more advanced Western models, a shortcut that would explain why the model's cyber skills plateau well below frontier performance.
How did Kimi K3 perform on exploit development?
The institutes used ExploitBench, a benchmark developed by Carnegie Mellon University that tracks how far a model advances through 41 real Chrome V8 engine vulnerabilities discovered after 2023. Each task moves through stages of the exploitation process, with Arbitrary Code Execution (ACE) sitting at the top. ACE gives attackers full control of a target system.
Kimi K3 never reached ACE on any of the 41 tasks. The leading US models hit that level on 20 of them. China's GLM-5.2 trailed further, scoring just 24.4 percent.

The institutes tested US closed-weight models with their system-level safeguards disabled to measure raw capability. Those safeguards remain active in public versions. Kimi K3, by contrast, didn't block exploit development or offensive operations at all. The model assisted with both without pushback.
What happened in the simulated network attack?
The second test, called "The Last Ones" (TLO), simulates a corporate network breach. It features a 32-step attack path across four subnets and roughly 20 hosts. A human expert would need about 20 hours to complete it.
Only a handful of models can solve TLO at all. The strongest US models complete it six or seven times out of ten attempts. Kimi K3 averaged step 17, meaning it got about halfway through. It finished the entire path once in ten tries while staying within the 100 million token limit. That single success shows the capability exists but can't be summoned reliably.

"Kimi K3 is capable of autonomously attacking small, weakly defended and vulnerable enterprise systems, when directed to do so and given initial network access," the British institute wrote. TLO doesn't simulate active defense, so it understates what real attackers would face. But even partial autonomous attack capability sets off alarms.
Are Chinese models catching up?
A time-series analysis by CAISI tracks cyber capabilities of US and Chinese models since early 2025 on an Elo-based scale. Both trend lines climb, but the Chinese curve consistently lags.

In a previous analysis, the British institute estimated the performance gap for open models at four to seven months, tighter than the six to ten months measured at the start of 2025. The new Kimi K3 results fit that pattern. Chinese open-weight models are improving, but they haven't closed the distance.
AISI cautioned against reading the gap as reassurance. Open models with growing cyber skills create "a persistent and irreversible risk of misuse" because anyone can download and run them without safeguards.
How agentic AI systems are being deployed in production environments
Why does distillation matter here?
The Kimi K3 findings lend weight to distillation allegations against Moonshot AI. US science advisor Michael Kratsios recently accused the company of training on outputs from Anthropic's Fable model, essentially using a stronger model's answers to teach a weaker one.
Distillation can transfer surface-level fluency and pattern recognition but often fails to capture deep reasoning on hard tasks. A model trained this way might sound capable on easy problems yet hit a ceiling on exploit chains or multi-step attacks. That's consistent with what the benchmarks show: Kimi K3 performs competently on simpler cyber tasks but stalls where the top US models keep climbing.
| Metric | Kimi K3 | GLM-5.2 | Leading US Models |
|---|---|---|---|
| ExploitBench Score | 32.2% | 24.4% | 76.2% |
| ACE Exploits (out of 41) | 0 | 0 | 20 |
| TLO Steps Completed (avg) | 17 | 11 | 28.5 |
| TLO Full Completion Rate | 1 in 10 | 0 in 10 | 6-7 in 10 |
| Safeguards Block Exploits? | No | Unknown | Yes (when enabled) |
Context on the infrastructure powering AI model training and deployment
What this means for model evaluation
The joint UK-US evaluation sets a template for how governments will assess frontier AI risk. ExploitBench and TLO test capabilities that matter for national security. The willingness of institutes to publish granular scores, and to test with safeguards disabled, signals that capability measurement is now a policy tool.
For teams building on open-weight models, the takeaway is sobering. Kimi K3 outperforms GLM-5.2 and sets a new benchmark among open-weight systems. But "best among open models" still means less than half the cyber capability of closed US systems. The gap matters when choosing which model to trust with sensitive infrastructure.
Logicity's Take
For AI builders evaluating Chinese open-weight models for production use, this report is a data point on capability ceilings. Kimi K3's inability to reach ACE on any exploit suggests the model lacks deep reasoning chains, not just training data. If distillation is the cause, expect similar ceilings in other open Chinese models built the same way. Teams should treat stated benchmark performance on reasoning tasks with skepticism until third-party evaluations confirm it. For red-teaming or security research, US closed models with safeguards disabled still set the capability bar.
Frequently Asked Questions
What is ExploitBench?
ExploitBench is a benchmark developed by Carnegie Mellon University that tests AI models on exploit development using 41 real Chrome V8 vulnerabilities discovered after 2023. It measures how far a model progresses through exploitation stages, with Arbitrary Code Execution as the highest level.
Did Kimi K3 refuse to help with cyber attacks?
No. The evaluation found that Kimi K3 assisted with exploit development and offensive cyber operations without pushback. Its safeguards did not block these requests.
How does Kimi K3 compare to other Chinese AI models?
Kimi K3 outperformed China's GLM-5.2, scoring 32.2% versus 24.4% on ExploitBench. It sets a new benchmark among open-weight models but trails leading US systems by a wide margin.
What is the distillation allegation against Moonshot AI?
US science advisor Michael Kratsios accused Moonshot AI of distilling Anthropic's Fable model, meaning they used Fable's outputs to train Kimi K3. This could explain why Kimi K3 performs reasonably on simpler tasks but fails on the hardest exploits.
Are Chinese AI models catching up to US models on cyber tasks?
Both are improving according to CAISI's Elo-based tracking, but Chinese models consistently trail. The gap has narrowed from six to ten months to four to seven months since early 2025.
Need Help Implementing This?
If you're evaluating open-weight models for security-sensitive applications or need help building AI safety assessments into your stack, reach out to Logicity's advisory team.
Source: The Decoder / Matthias Bastian
Manaal Khan
Tech & Innovation Writer
Produced with AI assistance and reviewed by the Logicity editorial team. Learn more in our Editorial Policy.
Related Articles
More in AI & Machine Learning
Bezos AI Lab Gets $10B: What Project Prometheus Means
Jeff Bezos is closing a $10 billion funding round for Project Prometheus, an AI lab focused on physics-based AI for manufacturing and engineering. With a $38 billion valuation and backing from JPMorgan and BlackRock, this signals a major shift in enterprise AI investment toward industrial applications.

Kimi K2.6 Open-Weight AI: 300 Agents at a Fraction of the Cost
Moonshot AI's Kimi K2.6 matches GPT-5.4 and Claude Opus 4.6 on coding benchmarks while running 300 parallel agents. For businesses locked into expensive API contracts, this open-weight model could slash AI infrastructure costs while delivering enterprise-grade automation.




