Every security scanner claims to find real bugs. The honest way to check is a test someone else wrote. So we run Vybscan on benchmarks built by others, with their answer keys, and compare against the scores other vendors publish.
Our latest test is the XBOW validation benchmarks: vulnerable web apps built as hacking challenges, with every planted bug labelled. We used the version that ZeroPath, an AI code scanner, prepared and published with its own scores. On ZeroPath's own test, Vybscan came out ahead on technical bugs and level on business logic.
The short version
- Vybscan found 82.9% of the technical vulnerabilities. ZeroPath found 80.0% on its own benchmark, Semgrep 57.1% and Snyk 40.0%.
- On business-logic bugs, such as one user opening another user's orders, rule-based tools find almost nothing. Vybscan's business-logic engine finds 87.5%, matching ZeroPath.
- We count only what the Vybscan dashboard shows as confirmed, after its own AI review. No raw noise is counted as a find.
- The whole test cost Vybscan about $13 in AI usage, including the fixed versions of the apps.
Why benchmarks built by others
A scanner vendor can always write a test it passes. A test built by someone else, with answer keys we did not write, is harder. We test against several: the XBOW validation benchmarks, RealVuln, OWASP Juice Shop and Cal.diy, the open-source edition of Cal.com. On Cal.diy Vybscan finds 91.7% of the known bugs, and on Juice Shop 85.0%.
Two kinds of bugs
Technical bugs are patterns in code: cross-site scripting, SQL injection, template injection, server-side request forgery. Business-logic bugs are missing rules: nothing checks that the order you open is yours. Rule-based scanners are built for the first kind and rarely catch the second. An AI scanner that reads the code can catch both.
Results
Technical vulnerabilities:
| Scanner | Detection |
|---|---|
| Vybscan | 82.9% |
| ZeroPath | 80.0% |
| Semgrep | 57.1% |
| Snyk | 40.0% |
| Bearer | 5.7% |
Business-logic vulnerabilities:
| Scanner | Detection |
|---|---|
| Vybscan | 87.5% |
| ZeroPath | 87.5% |
| Semgrep | 12.5% |
| Snyk | 0% |
| Bearer | 0% |
Other scanners' numbers are the scores ZeroPath published with the benchmark. Vybscan's are ours, scored as described below.
Better every week
Each benchmark also teaches us something. This one showed that Python apps load code in ways our analysis did not follow, so some Flask and Django routes looked unused. We fixed it for every customer. Across 61 other Python apps, code wrongly treated as unused fell from 150 files to 38.
How we scored it
- Dashboard view. A bug counts as found only if the Vybscan dashboard shows it as confirmed, after Vybscan's own AI review.
- Matching. A finding matches a label if it is in the labelled file, within 10 lines of the labelled range, and of the same vulnerability class. It is a fixed, repeatable rule.
- Scope. We scored the 43 labelled bugs whose code is in the published repository. The remaining labels sit in apps that were never ported.
Frequently asked questions
Why test on benchmarks built by others?
Because we did not write them. A test with someone else's answer key is the hardest place to look good, and the scores other vendors publish are the fairest comparison.
What is a business-logic vulnerability?
A missing rule rather than a bad pattern: a user can read someone else's records, set their own price, or claim a reward twice. These bugs are common in real apps and invisible to rule-based scanners.
Does Vybscan do more than code scanning?
Yes. It also scans dependencies, secrets and infrastructure settings on every pull request, for GitHub and GitLab. This benchmark covers only code scanning.
A note on these numbers: Vybscan changes every week. We expect these scores to move, and we will publish new results as they do.
Sources: XBOW validation benchmarks. ZeroPath's port of the benchmarks, with its published scores. RealVuln. MITRE CWE.
See what Vybscan finds in your code
Code, dependency and secrets scanning on every pull request, for GitHub and GitLab.
Try Vybscan →