Kimi K3 flunks cyber exam: only 40% as good as US models, distillation allegations surface
Two major AI safety institutions—the UK Artificial Intelligence Safety Institute (AISI) and the US Center for AI Standards and Innovation (CAISI)—put Moonshot's latest model, Kimi K3, through a rigorous security exam. The results? Not pretty.
ExploitBench: A Tale of Two Scores
The first test, Carnegie Mellon University's ExploitBench, presented 41 real vulnerabilities from the Chrome V8 engine discovered after 2023. Models were scored on how far they could progress in exploiting each flaw. Leading US models averaged 76.2%. Kimi K3? Just 32.2%. China's Zhipu GLM-5.2 trailed at 24.4%.
More telling: US models achieved "arbitrary code execution" (ACE)—the most dangerous level, giving attackers full control—in 20 out of 41 tasks. Kimi K3 didn't reach ACE in a single one. Notably, the testers disabled system-level protections on US closed-weight models to measure their ceiling, while those protections remain on by default in public versions.

The Last Survivor: Stuck at Step 17
The second test, called "The Last Survivor" (TLO), simulated a real enterprise network attack spanning four subnets and about 20 hosts, with 32 steps. Human experts need about 20 hours to complete it. Only a handful of models have passed; the strongest US model succeeds 6-7 times out of 10.
Kimi K3 typically stalled at step 17. US models reached step 28.5 on average. GLM-5.2 only made it to step 11. But here's a twist: in one of ten attempts, Kimi K3 completed the entire attack chain within 1 billion tokens. So the capability exists—it's just not consistent. The evaluators concluded that, given initial access and an extra instruction, Kimi K3 could independently conquer smaller, less defended enterprise systems. (TLO lacks active defenses, so it's not fully realistic.)
The Bigger Picture: Catching Up, But Still Behind
CAISI has tracked cyber capabilities of Chinese and US models since early 2025 using an Elo rating system. Both trend lines rise, but a persistent gap remains. A 400-point Elo increase means a tenfold boost in task completion probability—so the gap is significant. Earlier, AISI estimated open-source models lag US systems by 4-7 months; at the start of 2025, it was 6-10 months. The latest results fit that pattern.
AISI's warning is blunt: the growing cyber capabilities of open-source models pose a "continuous and irreversible misuse risk."
The Distillation Connection
The most intriguing part is how these test results intersect with recent accusations. US science advisor Michael Krazios publicly claimed that Moonshot used outputs from Anthropic's Fable model (Claude) to "distill" and improve Kimi K3, and that it obtained US-controlled NVIDIA GB300 chips.
The pattern fits: Kimi K3 scores well on general benchmarks but bombs on cybersecurity. If it mainly learned from Claude's outputs on general knowledge, programming, and agent tasks, that would explain the gap. Anthropic's safety classifier filters out advanced cyber queries, so those samples were scarce in the distilled dataset. The model could compete on standard rankings but didn't learn deeper exploit techniques.
AISI's evaluation indirectly supports this: disabling system-level protections on US models revealed cyber capabilities nearly impossible to access through public interfaces—and thus unlikely to appear in distilled data. One exam, two revelations: a score, and the unspoken origins beneath it.
Key Points:
- Kimi K3 scored 32.2% on ExploitBench vs. 76.2% for US models; never achieved arbitrary code execution.
- In simulated enterprise attack (TLO), Kimi K3 averaged step 17 of 32; US models reached step 28.5.
- Chinese models are improving but still lag US counterparts by 4-10 months.
- Poor cyber performance aligns with distillation allegations: Kimi K3 likely trained on Claude outputs that lack advanced exploit data.
- AISI warns open-source cyber capabilities pose continuous misuse risk.