VulnBench: Can LLMs find the same security bugs twice?

lirantal1 pts0 comments

Snyk VulnBenchSkip to contentMenu

Current releaseSnyk VulnBench JS 1.0<br>A Snyk benchmark initiative<br>Can LLMs find the same bugs twice?<br>A repeatability and Snyk-reference agreement study<br>We ran the same agentic security review five times against inspectable JavaScript projects to measure what recurs, what varies, and how model findings align with a deterministic Snyk Code reference set.<br>Explore the resultsRead the paper<br>View methodologyDownload data<br>300scans<br>10projects<br>6configurations<br>5repetitions

Headline evidence · 5 identical reviews<br>Same review. Different results.

JS 1.0<br>300 scans · 10projects · 6 configurations<br>Reference-matched findings seen in all five runs84.8%

Unmatched findings seen in all five runs13.7%

Unmatched findings seen in only one run49.7%

134 of 158 Reference-matched findings seen in all five runsInspect<br>22 of 161 Unmatched findings seen in all five runsInspect<br>80 of 161 Unmatched findings seen in only one runInspect<br>For AppSec teams One run is not a decision. Repeatability helps prioritize what to inspect; it does not label unmatched reports as false positives.<br>134 of 158 · Reference-matched findings seen in all five runs<br>View exact recurrence valuesFinding groupCountShareReference-matched findings seen in all five runs134 of 15884.8%Unmatched findings seen in all five runs22 of 16113.7%Unmatched findings seen in only one run80 of 16149.7%Dataset 1.0.0Snyk VulnBench JS 1.0Unique signatures · recurrence across 5 identical runsSource: published JS 1.0 paper

Latest evidence · JS 1.0<br>Same code. Same prompt. Different findings.<br>Across five identical runs, 84.8% of reference-matched findings recurred every time. Nearly half of unmatched reports appeared only once—evidence to inspect, not dismiss.

5.67×Opus 4.7 Max cost more and scored lower than Opus 4.6 Medium: 68.8% vs 75.4% Snyk-reference F1.Inspect evidence →49.7%Nearly half of unmatched reports surfaced in just one run.Inspect evidence →13.7%Only 13.7% of unmatched reports recurred in all five runs.Inspect evidence →

What the results mean in practice<br>An AI review is a measurement—not a verdict.<br>Repeat the same task and the story sharpens: many reference-matched findings hold steady; AI review and deterministic SAST reveal different blind spots; and more spend does not reliably improve Snyk-reference agreement.

01Repeatability makes confidence visible<br>134 of 158 reference-matched findings appeared in every one of five identical reviews. Repetition shows which reported patterns persist instead of relying on a single run.<br>Interpret with care: A reference match measures agreement with Snyk Code, not independent ground-truth accuracy.

Open the evidence<br>02AI review and SAST expose different blind spots<br>Models surfaced high-signal exploit shapes, while deterministic Snyk Code consistently enumerated repeated data-flow sinks. Their differing results are useful evidence—not a reason to declare one a universal winner.<br>Interpret with care: Unmatched findings require case-level inspection before they can be classified.

Open the evidence<br>03Paying more did not reliably improve the result<br>In this benchmark, higher session cost did not consistently yield higher Snyk-reference F1. Spend alone is a poor shortcut for choosing a configuration.<br>Interpret with care: Cost estimates reflect the tested small fixtures and publication assumptions.

Open the evidence

Benchmark anatomy<br>How VulnBench measures behavior<br>Repeat the conditions, preserve the evidence, and separate observed agreement from claims the protocol cannot support.<br>Read the full methodology<br>1Select inspectable projects<br>Ten small JavaScript and Express fixtures make every run and reference finding reviewable.

2Repeat the same task<br>Each configuration sees the same code, prompt, harness, and task five times.

3Normalize findings<br>Reported issues become documented signatures suitable for recurrence analysis.

4Match the reference set<br>The scorer compares vulnerability type against deterministic Snyk Code findings.

5Measure behavior<br>Agreement, recurrence, variance, coverage, cost, tokens, and duration stay distinct.

6Inspect divergence<br>Unmatched reports remain evidence to investigate—not automatic false positives.

Research principles<br>Evidence before ranking<br>VulnBench is a versioned research initiative, not a universal leaderboard. Every release defines what it measured and what it did not prove.

Transparent reference setsDefinitions and limitations stay visible.<br>Repeated measurementVariance is a result, not a footnote.<br>Inspectable casesHeadline claims link toward underlying evidence.<br>Reproducible dataVersioned source artifacts remain downloadable.<br>Explicit limitationsAgreement is never relabeled as accuracy.<br>Release history<br>Snyk VulnBench JS 1.0 Published 11 June 2026 · Current<br>Further releases are planned. Unpublished results will not appear as speculative rankings or empty release cards.<br>View release catalog

Publication<br>Read, reproduce, and cite the work<br>The paper, reviewed...

findings reference evidence five unmatched snyk

Related Articles