A Snyk benchmark initiative
Can LLMs find the same bugs twice?
A repeatability and Snyk-reference agreement study
We ran the same agentic security review five times against inspectable JavaScript projects to measure what recurs, what varies, and how model findings align with a deterministic Snyk Code reference set.
- 300
- scans
- 10
- projects
- 6
- configurations
- 5
- repetitions
Headline evidence
Stable matches. Variable extras.
- 134 of 158Reference-matched findings seen in all five runsInspect
- 80 of 161Unmatched findings seen in only one runInspect
- 22 of 161Unmatched findings seen in all five runsInspect
134 of 158 · Reference-matched findings seen in all five runs
View exact recurrence values
| Finding group | Count | Share |
|---|---|---|
| Reference-matched findings seen in all five runs | 134 of 158 | 84.8% |
| Unmatched findings seen in only one run | 80 of 161 | 49.7% |
| Unmatched findings seen in all five runs | 22 of 161 | 13.7% |
Latest evidence · JS 1.0
Repeatability separated signal from noise.
Reference-matched findings were markedly more stable across identical runs. Unmatched does not mean false positive; it identifies evidence that needs inspection.
Observed results
What the data show
Three findings define this release. Each one links to the measured view and keeps its interpretation boundary attached.
- 01Open the evidence
Reference-matched findings were substantially more repeatable
134 of 158 unique reference-matched findings appeared in all five repetitions.
Interpret with care: A reference match measures agreement with Snyk Code, not independent ground-truth accuracy.
- 02Open the evidence
LLM review and deterministic SAST showed complementary behavior
Models identified high-signal exploit shapes while deterministic SAST systematically enumerated repeated data-flow sinks.
Interpret with care: Unmatched findings require case-level inspection before they can be classified.
- 03Open the evidence
Higher session cost did not guarantee higher reference agreement
The published cost-quality comparison does not show a monotonic relationship between spend and Snyk-reference F1.
Interpret with care: Cost estimates reflect the tested small fixtures and publication assumptions.
Benchmark anatomy
How VulnBench measures behavior
Repeat the conditions, preserve the evidence, and separate observed agreement from claims the protocol cannot support.
Read the full methodology- 1
Select inspectable projects
Ten small JavaScript and Express fixtures make every run and reference finding reviewable.
- 2
Repeat the same task
Each configuration sees the same code, prompt, harness, and task five times.
- 3
Normalize findings
Reported issues become documented signatures suitable for recurrence analysis.
- 4
Match the reference set
The scorer compares vulnerability type against deterministic Snyk Code findings.
- 5
Measure behavior
Agreement, recurrence, variance, coverage, cost, tokens, and duration stay distinct.
- 6
Inspect divergence
Unmatched reports remain evidence to investigate—not automatic false positives.
Research principles
Evidence before ranking
VulnBench is a versioned research initiative, not a universal leaderboard. Every release defines what it measured and what it did not prove.
- Transparent reference sets
- Definitions and limitations stay visible.
- Repeated measurement
- Variance is a result, not a footnote.
- Inspectable cases
- Headline claims link toward underlying evidence.
- Reproducible data
- Versioned source artifacts remain downloadable.
- Explicit limitations
- Agreement is never relabeled as accuracy.
Release history
Further releases are planned. Unpublished results will not appear as speculative rankings or empty release cards.
View release catalogPublication
Read, reproduce, and cite the work
The paper, reviewed methodology, source snapshot, and release data use stable public links.
Preferred citation
Liran Tal, Johannes Kloos, Arsenii Rudich, Stephen Thoemmes, Manoj Nair. “Snyk VulnBench JS 1.0: Can LLMs Find the Same Bugs Twice?.” arXiv, 2026.