How Far Behind the Frontier Are Leading Open Weight Models on Cyber?

herbertl1 pts0 comments

How Far Behind the Frontier are Leading Open Weight Models on Cyber? | AISI Work

Read the Frontier AI Trends Report<br>Please enable javascript for this website.

Careers

Blog

Cyber & Autonomous Systems

How Far Behind the Frontier are Leading Open Weight Models on Cyber?

We evaluated the cyber capabilities of leading open and closed weight AI models, and found that recent open models GLM-5.2 and DeepSeek V4-Pro perform similarly to frontier closed models released 4 to 7 months before them – a narrower gap than the 6 to 10 months we measured through most of 2025.

Jul 17, 2026

AISI has tracked the cyber capabilities of frontier AI models since 2023. On our evaluations, the most capable models have consistently been closed weight – models whose parameters are private, accessible only through developer-controlled interfaces. Leading open weight models – whose parameters anyone can download, run, and modify – have consistently trailed behind. How far behind, whether the gap is closing, and how quickly, are all live questions for researchers, policymakers, and cyber defenders.<br>Open weight models bring real benefits. They can be hosted privately, with no data returning to the model providers, adapted to specific tasks, and run only at the cost of compute. Once in use, they offer a dependable base that providers can’t change or deprecate. They enable open collaboration and innovation, as well as certain types of safety research requiring access to model weights – including some of the work of AISI’s Model Transparency team.<br>But they can also carry risk. The same openness underpinning these benefits precludes many of the safety measures that closed model developers can use to detect and disrupt misuse, iterate on safeguards as vulnerabilities emerge, control user access and withdraw models. Once open weight models are released, these options are lost permanently: safeguards can be removed, and copies can be downloaded, redistributed, and run on private systems beyond monitoring. For models with dangerous capabilities – including highly cyber-capable models – open weight release therefore creates a persistent and irreversible risk of misuse.<br>One reason the gap between the cyber capabilities of open and closed models matters is because it provides a preparation time: a window for cyber defenders with access to the most capable closed systems to take action before today’s frontier cyber capabilities might become available without the same safeguards. This becomes more pressing as frontier AI cyber capabilities continue to advance; in April 2026, two closed models, Mythos Preview and GPT-5.5, demonstrated some of the largest jumps in AI cyber capability AISI has observed since testing began, prompting international warnings to act on a rapidly transforming cyber risk landscape.<br>This is our first public analysis of how far leading open weight models trail the closed cyber frontier. Our evaluations find that GLM-5.2 (June 2026) was the most cyber-capable open weight model at time of testing. It performs similarly to Opus 4.6 (Feb 2026) on AISI’s narrow cyber tasks and Opus 4.5 (Nov 2025) on our longer-horizon cyber ranges, meaning it trails the frontier by 4 to 7 months. This is narrower than the 6 to 10 month gap we measured in internal evaluations of open weight models released from January to September 2025.<br>Below, we lay out these results, which models we tested, and what a narrowing gap means for cyber defence.<br>Results: the current cyber gap for open weight AI<br>Our evaluations take a broad view of models’ cyber capabilities. AISI’s narrow cyber tasks assess specific cyber skills across four difficulty levels. Separately, our cyber ranges assess autonomous cyber capability – a model’s ability to sustain end-to-end planning and execution over long-horizon, multi-step cyberattacks in simulated networks containing vulnerabilities.<br>Here we share results for two open weight models we tested, selected due to their candidacy to lead open weight cyber capability at the time of release, GLM-5.2 and DeepSeek V4-Pro. We also cite results from similar, internal testing conducted in 2025. AISI intends to test Kimi K3 on this same basis, once its weights are publicly released (announced for end of July).<br>Narrow cyber tasks<br>AISI’s narrow cyber tasks measure the difficulty of tasks a model can complete, ranging from “technical non-expert” (novices with some technical expertise but limited cyber knowledge, e.g. a data analyst) to “expert” (tasks typically requiring deep knowledge of cybersecurity).1 They span several cybersecurity capabilities such as vulnerability research and exploitation, reverse engineering, web exploitation, and cryptography.

Figure 1: Average success rate on 70 of AISI’s narrow cyber tasks, given 5 attempts per task with a 2.5M token limit per attempt. 2 Dotted lines connect open and closed model with comparable performance, indicating the gap between release dates.On these tasks, GLM-5.2 performs...

cyber models open weight frontier aisi

Related Articles