UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities | NIST
Skip to main content
Official websites use .gov
A .gov website belongs to an official government organization in the United States.
Secure .gov websites use HTTPS
A lock (
Lock<br>A locked padlock
) or https:// means you’ve safely connected to the .gov website. Share sensitive information only on official, secure websites.
https://www.nist.gov/news-events/news/2026/07/uk-aisi-caisi-preliminary-assessment-kimi-k3s-cyber-capabilities
UPDATES
UK AISI / CAISI Preliminary Assessment of Kimi K3's Cyber Capabilities
July 23, 2026
Share
X.com
The UK Artificial Intelligence Security Institute (UK AISI) and the U.S. Center for AI Standards and Innovation (CAISI) (UK AISI / CAISI) conducted a joint evaluation of Moonshot AI’s latest model, Kimi K3 (released on July 16, 2026 and slated for open-weight release by July 27, 2026). This evaluation focused on Kimi K3's cyber capabilities and found that:<br>Kimi K3 performs significantly below the most recent frontier cyber-capable models on preliminary cyber evaluations run by UK AISI / CAISI. Specifically:When tasked to develop exploits, Kimi K3 performs significantly below the most recent frontier cyber-capable models (Figure 1).<br>When tasked to attack a simulated corporate network (“The Last Ones”), Kimi K3 performs significantly below the most recent frontier cyber-capable models. Specifically, on average, Kimi K3 reached step 17 of this 32-step attack path, while the most cyber-capable U.S. models reached 28.5 steps on average.
Kimi K3 performs above GLM-5.2 on the same preliminary cyber evaluations run by UK AISI / CAISI (Figure 2).<br>Kimi K3’s safeguards allow assistance with agentic cyber exploit development . Kimi K3's safeguards did not prevent it from attempting cyber exploit development or offensive cyber operations during UK AISI / CAISI's evaluations.
Exploit Development Capability
Credit:
CAISI/NIST
Figure 1: Performance of Kimi K3 and other models on an exploit development benchmark (ExploitBench) . Higher success rate indicates greater cyber capability. Error bars represent 95% confidence intervals. ExploitBench measures the capability of a model to develop end-to-end exploits given a vulnerability.
Detailed Results<br>These results represent preliminary evaluations on a small set of public and private benchmarks. U.S. closed-weight models were evaluated with system-level safeguards disabled to reduce refusals and enable measurement of maximal capabilities. Publicly available versions of these models have these safeguards enabled. Due to the specifics of Kimi K3’s hosting setup, UK AISI / CAISI ran a selective set of cyber evaluations. Detailed methodologies are provided in individual sections.<br>Cyber Capability Trends<br>The cyber capability of models is aggregated across multiple tasks from multiple benchmarks using an approach inspired by Item Response Theory (IRT). For details of the methodology please see prior published reports. Kimi K3’s overall cyber capability has a larger confidence interval than other models because it was estimated from a single benchmark (ExploitBench, which has 41 tasks focused on exploit development). ExploitBench is a leading benchmark to measure a model’s ability to progress along the software exploitation ladder. All other models’ overall cyber capability scores were derived from a larger number of tasks that covered additional domains of cyber capability.
Credit:
CAISI/NIST
Figure 2: Preliminary comparison of aggregate capabilities over time of the most capable U.S. and PRC models as of Kimi K3’s release. The U.S. trendline is composed of results from frontier U.S. models . A 400-point increase on the y-axis equates to a 10x increase in the odds of solving tasks. Error bars and shaded regions denote 95% CIs.
Exploit Development: ExploitBench<br>ExploitBench is a public benchmark, developed by Carnegie Mellon University, that measures a model’s ability to progress along the software exploitation ladder, including coverage and crash reproduction, arbitrary read/write, control flow hijack, and arbitrary code execution. The benchmark tests models on 41 recent (post-2023) vulnerabilities in the V8 engine (the JavaScript and WebAssembly software that powers Chrome).<br>ExploitBench results are presented in Figures 1 and 3.
Credit:
CAISI/NIST
Figure 3: Detailed ExploitBench performance for Kimi K3 and other models . Darker shading indicates greater cyber capability. Each row represents a key milestone in the exploit development chain, and each cell shows the number of ExploitBench tasks for which the model(s) in question were able to reach that milestone.
Kimi K3 outperforms GLM-5.2, the most cyber-capable open-weight model as of June 2026 . Kimi K3 achieves a score of 32%, whereas GLM-5.2 achieves a score of 24% (Figure 1).<br>Unlike the most cyber-capable models, Kimi K3 failed to develop exploits that achieved...