JS 1.0 · Interactive explorer
Test the published conclusions against the evidence
Filter configurations and projects, move between guided analytical views, compare up to four configurations, and share deterministic explorer state.
Snyk VulnBench JS 1.0
JS 1.0 explorer
- Dataset
- 1.0.0
- View
- Summary
- Filters
- 0 active filters
- Records
- 300 represented runs
- Metric
- f1
- Aggregation
- mean
Published default view
Configuration summary
Compare Snyk-reference agreement, repeated-run spread, coverage, and resource use under the active project and configuration filters.
| Configuration | |||||||
|---|---|---|---|---|---|---|---|
| Snyk Code SASTDeterministic reference reproduction | 100.0% | 0.0 pp | 100.0% | 100.0% | 14.8 s | 0 | N/A |
| Claude Opus 4.6 MediumModel via agentic harness | 75.4% | 0.2 pp | 68.0% | 91.5% | 27.3 s | 51,574 | $0.063 |
| Claude Opus 4.6 HighModel via agentic harness | 75.2% | 0.3 pp | 68.2% | 89.8% | 53.8 s | 66,929 | $0.125 |
| Claude Opus 4.7 MaxModel via agentic harness | 68.8% | 2.2 pp | 71.4% | 69.6% | 37.4 s | 95,969 | $0.356 |
| Claude Sonnet 4.6 MediumModel via agentic harness | 67.4% | 0.9 pp | 80.9% | 62.6% | 59.3 s | 56,992 | $0.086 |
| Claude Sonnet 4.6 HighModel via agentic harness | 64.9% | 3.5 pp | 81.3% | 58.6% | 94.8 s | 74,240 | $0.132 |
Macro average across 10 fixtures and 5 repetitions · Dataset 1.0.0 · Snyk VulnBench JS 1.0
Interpretation boundary: Snyk Code’s 100% row is deterministic reproduction of the reference set it defines. It is not a universal accuracy result.
Snyk Code SASTF1 standard deviation: 0.0 ppSnyk-reference F1: 100.0%
| Configuration | F1 standard deviation | Snyk-reference F1 |
|---|---|---|
| Snyk Code SAST | 0.0 pp | 100.0% |
| Claude Opus 4.6 Medium | 0.2 pp | 75.4% |
| Claude Opus 4.6 High | 0.3 pp | 75.2% |
| Claude Opus 4.7 Max | 2.2 pp | 68.8% |
| Claude Sonnet 4.6 Medium | 0.9 pp | 67.4% |
| Claude Sonnet 4.6 High | 3.5 pp | 64.9% |
Upper-left means higher Snyk-reference agreement and lower repeated-run variance.
Caveat: Reference agreement is not universal vulnerability-detection accuracy.
Snyk VulnBench JS 1.0 · Dataset 1.0.0 · Macro average across 10 fixtures and 5 repetitions · Source /data/js-1.0/published-evidence.json
Headline recurrence
Stable reference matches contrast with variable unmatched reports across identical runs. This release-wide context is not changed by active project or configuration filters.
Claude Opus 4.6 MediumEstimated model-session cost: $0.063Snyk-reference F1: 75.4%
| Configuration | Estimated model-session cost | Snyk-reference F1 |
|---|---|---|
| Claude Opus 4.6 Medium | $0.063 | 75.4% |
| Claude Opus 4.6 High | $0.125 | 75.2% |
| Claude Opus 4.7 Max | $0.356 | 68.8% |
| Claude Sonnet 4.6 Medium | $0.086 | 67.4% |
| Claude Sonnet 4.6 High | $0.132 | 64.9% |
Upper-left means higher Snyk-reference agreement at lower estimated cost.
Caveat: Costs reflect small fixtures and published model-session assumptions.
Snyk VulnBench JS 1.0 · Dataset 1.0.0 · Macro average across 10 fixtures and 5 repetitions · Source /data/js-1.0/published-evidence.json
Compare
0 pinned
Pin up to four configurations from a view to compare agreement, repeatability, and resource use.