Current releaseSnyk VulnBench JS 1.0

A Snyk benchmark initiative

Can LLMs find the same bugs twice?

A repeatability and Snyk-reference agreement study

We ran the same agentic security review five times against inspectable JavaScript projects to measure what recurs, what varies, and how model findings align with a deterministic Snyk Code reference set.

300
scans
10
projects
6
configurations
5
repetitions

Headline evidence

Stable matches. Variable extras.

JS 1.0

134 of 158 · Reference-matched findings seen in all five runs

View exact recurrence values
Finding groupCountShare
Reference-matched findings seen in all five runs134 of 15884.8%
Unmatched findings seen in only one run80 of 16149.7%
Unmatched findings seen in all five runs22 of 16113.7%
Dataset 1.0.0Snyk VulnBench JS 1.0Unique signatures · recurrence across 5 identical runs

Latest evidence · JS 1.0

Repeatability separated signal from noise.

Reference-matched findings were markedly more stable across identical runs. Unmatched does not mean false positive; it identifies evidence that needs inspection.

Observed results

What the data show

Three findings define this release. Each one links to the measured view and keeps its interpretation boundary attached.

  1. 01

    Reference-matched findings were substantially more repeatable

    134 of 158 unique reference-matched findings appeared in all five repetitions.

    Interpret with care: A reference match measures agreement with Snyk Code, not independent ground-truth accuracy.

    Open the evidence
  2. 02

    LLM review and deterministic SAST showed complementary behavior

    Models identified high-signal exploit shapes while deterministic SAST systematically enumerated repeated data-flow sinks.

    Interpret with care: Unmatched findings require case-level inspection before they can be classified.

    Open the evidence
  3. 03

    Higher session cost did not guarantee higher reference agreement

    The published cost-quality comparison does not show a monotonic relationship between spend and Snyk-reference F1.

    Interpret with care: Cost estimates reflect the tested small fixtures and publication assumptions.

    Open the evidence

Benchmark anatomy

How VulnBench measures behavior

Repeat the conditions, preserve the evidence, and separate observed agreement from claims the protocol cannot support.

Read the full methodology
  1. 1

    Select inspectable projects

    Ten small JavaScript and Express fixtures make every run and reference finding reviewable.

  2. 2

    Repeat the same task

    Each configuration sees the same code, prompt, harness, and task five times.

  3. 3

    Normalize findings

    Reported issues become documented signatures suitable for recurrence analysis.

  4. 4

    Match the reference set

    The scorer compares vulnerability type against deterministic Snyk Code findings.

  5. 5

    Measure behavior

    Agreement, recurrence, variance, coverage, cost, tokens, and duration stay distinct.

  6. 6

    Inspect divergence

    Unmatched reports remain evidence to investigate—not automatic false positives.

Research principles

Evidence before ranking

VulnBench is a versioned research initiative, not a universal leaderboard. Every release defines what it measured and what it did not prove.

Transparent reference sets
Definitions and limitations stay visible.
Repeated measurement
Variance is a result, not a footnote.
Inspectable cases
Headline claims link toward underlying evidence.
Reproducible data
Versioned source artifacts remain downloadable.
Explicit limitations
Agreement is never relabeled as accuracy.

Release history

Snyk VulnBench JS 1.0Published 11 June 2026 · Current

Further releases are planned. Unpublished results will not appear as speculative rankings or empty release cards.

View release catalog

Publication

Read, reproduce, and cite the work

The paper, reviewed methodology, source snapshot, and release data use stable public links.

Preferred citation

Liran Tal, Johannes Kloos, Arsenii Rudich, Stephen Thoemmes, Manoj Nair. “Snyk VulnBench JS 1.0: Can LLMs Find the Same Bugs Twice?.” arXiv, 2026.

Liran Tal · Johannes Kloos · Arsenii Rudich · Stephen Thoemmes · Manoj Nair