About VulnBench
A durable home for inspectable benchmark evidence
Snyk VulnBench publishes versioned studies of vulnerability-finding behavior, their methods, datasets, cases, and interpretation boundaries.
Product promise
Explore the result. Inspect the evidence.
Explore how reliably AI systems find vulnerabilities—and inspect the evidence behind every conclusion.
VulnBench is not a universal leaderboard, a live scanner, or the home for all Snyk AI security research. It is one benchmark initiative spanning versioned protocols and stable public releases.
Initiative principles
Scientific framing stays in the interface
- Evidence before ranking. Lead with measured behavior.
- Precise language. Reference agreement is not renamed accuracy.
- Every value has provenance. Release, dataset, metric, aggregation, and sample size travel together.
- Version boundaries are explicit. Incompatible protocols do not produce false comparisons.
- Immutable releases. Corrections are documented rather than silently rewritten.
Team
Researchers and authors
- Liran TalSnyk
- Johannes KloosSnyk
- Arsenii RudichSnyk
- Stephen ThoemmesSnyk
- Manoj NairSnyk
Citation guidance
Cite the release, not an unversioned score
Liran Tal, Johannes Kloos, Arsenii Rudich, Stephen Thoemmes, Manoj Nair. “Snyk VulnBench JS 1.0: Can LLMs Find the Same Bugs Twice?.” arXiv, 2026.
Include the release name and dataset version1.0.0 when reproducing a chart or filtered result.
Open the canonical paper