Research and findings
Citable findings, benchmarks, and original analysis from the DevSecCode Team. Each entry includes a headline number, methodology, and a ready-to-paste citation block.
AI Code Security
About four in five functionally correct solutions from the evaluated AI coding agent were insecure.
~79% of functionally correct agent solutions were insecure
Summary
The SusVibes Benchmark tested 200 coding tasks across 77 CWE types against AI coding agents. In the first version of the paper (December 2025), SWE-Agent with Claude 4 Sonnet produced functionally correct solutions for 61% of tasks, but only 10.5% of tasks were solved securely, so 82.8% of its functionally correct solutions failed the security tests. The current version of the paper reports that about 79% of functionally correct solutions were insecure, with vulnerabilities including injection, authentication bypass, IDOR, and exposed network services.
Methodology
Each task pairs a benign functional specification with adversarial security test cases. A solution counts as correct only if it passes the functional spec; it counts as secure only if it also passes the security tests. The headline figure is the percentage of functionally correct solutions that fail one or more security tests: 82.8% for SWE-Agent with Claude 4 Sonnet in the first version, and about 79% in the current version.
Source
SusVibes Benchmark · CMU, Columbia, Johns Hopkins (December 2025)
Suggested citation
SusVibes Benchmark (arXiv:2512.03262) via DevSecCode, "AI-generated code security: about four in five functionally correct solutions from the evaluated agent were insecure" (devseccode.com/articles/research, 2026).Real-World Disaster
The OpenClaw incident exposed 30,000+ systems with AI-generated security holes.
30,000+ systems compromised via AI-generated code
Summary
OpenClaw was a viral AI assistant project that reached 100,000 GitHub stars in two months. Subsequent security analysis found that AI-generated portions of the codebase contained catastrophic vulnerabilities that traditional scanners missed: a single-character password ('a') was accepted as valid authentication, services bound to 0.0.0.0:18789 exposed administrative interfaces to the public internet, the AI freely returned API keys when asked, and an allowInsecureAuth: true flag bypassed all authentication checks. By the time the issues were disclosed, 30,000+ deployed instances had been identified by external scanners.
Methodology
Population of affected systems estimated via Shodan/Censys scans for the distinctive OpenClaw service banner on port 18789 during the January 2026 disclosure window. The vulnerability classes (CWE-521 weak authentication, CWE-668 exposed services, CWE-266 security bypass, CWE-200 credential exposure, CWE-78 command injection) are deterministically detectable at code level.
Source
Bitsight Security Research (January 2026)
Suggested citation
Bitsight Security Research via DevSecCode, "OpenClaw incident: 30,000+ systems exposed by AI-generated security holes" (devseccode.com/articles/research, 2026).Deva Coder Benchmark
Deva Coder v8 achieves 87.5% accuracy on SecurityEval CWE detection.
87.5% SecurityEval CWE detection accuracy
Summary
Deva Coder v8, the local security-focused coding model in the Deva model family, was benchmarked against the SecurityEval CWE detection suite. The model achieved 87.5% accuracy on classifying and remediating CWE-categorized vulnerabilities in code, 99.7% syntax pass rate on MBPP, 93.3% tool-use compliance, and 100% fix generation rate when a vulnerability is identified. The model runs locally on Apple Silicon or H200-class GPUs with no cloud calls.
Methodology
SecurityEval is an open benchmark covering ~75 CWE patterns across Python, JavaScript, TypeScript, Go, Java, and Ruby. Each task supplies vulnerable source code; the model must identify the CWE and produce a remediation. Accuracy is the percentage of tasks where the model correctly identifies the CWE and produces a remediation that passes the security test suite. First-token latency was ~1.3s on H200; locally on Apple Silicon, the model produces secure code without any outbound network calls.
Source
Deva Coder v8 benchmark · H200 GPU run, April 2026
Suggested citation
DevSecCode, "Deva Coder v8 SecurityEval results: 87.5% CWE detection accuracy" (devseccode.com/articles/research, 2026).Supply Chain Surface
CVE advisories cross-referenced against a 2,800+ package metadata catalog.
2,800+ packages in supply chain catalog
Summary
Deva's SCA layer maintains a CVE advisory catalog synced from the National Vulnerability Database, the GitHub Advisory Database, and the Open Source Vulnerabilities database. The catalog is enriched with metadata for 2,800+ packages across npm, PyPI, RubyGems, Maven Central, Go modules, and Crates. SCA runs locally without contacting external services in air-gapped deployments by using a periodically-refreshed snapshot of the catalog.
Methodology
CVE advisories are de-duplicated across NVD, GHSA, and OSV using purl (package URL) identifiers. Package metadata (download counts, last-published date, maintainer counts, repository linkage) is sourced from native registry APIs.
Source
Deva supply-chain catalog · synced from NVD, GHSA, OSV
Suggested citation
DevSecCode, "Deva supply-chain catalog: CVE advisories across 2,800+ packages" (devseccode.com/articles/research, 2026).Compliance Coverage
17 compliance frameworks mapped at code level with 6 report export formats.
17 compliance frameworks mapped
Summary
Deva's compliance engine maps every CWE finding to the relevant controls of 17 compliance frameworks: HIPAA, PCI-DSS v4.0.1, SOC 2 Type II, CMMC (Levels 1 through 3), NIST SP 800-53 Rev 5, NIST CSF 2.0, FedRAMP Rev 5 (Low, Moderate, High), GDPR, ISO 27001:2022, ISO/IEC 27701, OWASP Top 10 (2025), CIS Controls v8, HITRUST CSF v11, CCPA/CPRA, DORA, NIS2, and the EU Cyber Resilience Act. Findings export in SARIF, JUnit XML, OSCAL Assessment Results, Markdown, signed evidence packages, and an agent-json format consumable by downstream AI agents. Compliance results also export as OSCAL, POA&M, and SPRS.
Methodology
Each compliance framework's controls are mapped to specific CWE rules via a many-to-many mapping table maintained by the DevSecCode Team. The mapping is bidirectional: a finding shows which controls it violates, and a framework view shows which controls have passing, failing, or attestation-required status. SARIF and OSCAL exports include the compliance metadata so downstream tools (audit evidence platforms, SIEMs, GRC systems) can consume the data directly.
Source
Deva compliance engine · 17 frameworks shipping as of 2026-05
Suggested citation
DevSecCode, "Deva compliance engine: 17 frameworks mapped at code level" (devseccode.com/articles/research, 2026).Citation policy: Findings on this page are intended for use as references in academic, industry, and journalistic work. Each item lists its source and a suggested citation string. If a finding cites an external source (SusVibes Benchmark, Bitsight Security Research), follow that source's own citation policy in addition. Direct anchor links work for each finding (for example, /articles/research#susvibes-ai-code-insecure).