# Research: ~90% of Leaked Malware Source Code Contains Exploitable Software Weaknesses (Vouvoutsis, Patsakis & Casino, arXiv:2606.05945)

> Academic researchers static-analyzed 658 leaked malware source-code projects from VX-Underground (paired against 249 benign open-source projects) with Cppcheck, Bandit, Snyk, and Semgrep and found close to 90% contained at least one recognized software weakness. Disabled TLS certificate validation in C2 clients (CWE-295) leaves command-and-control traffic open to interception, and over 40% of weaknesses were shared across multiple malware families (e.g. WannaCry, Emotet, Torpig) — meaning defenders can turn bugs in malware itself into a disruption opportunity.

- **Published:** 2026-06-09T00:00:00Z
- **Last reviewed:** 2026-06-09T00:00:00Z
- **Canonical:** https://intel.threadlinqs.com/threat/TL-2026-0739
- **ID:** TL-2026-0739
- **Severity:** INFO
- **Category:** THREAT_INTEL
- **Status:** ACTIVE
- **Detections:** 9 · **IOCs:** 26 (full data via the Threadlinqs MCP server — Purple tier)

## Description

This is a defensive threat-intelligence research note, not a product vulnerability. Vasilis Vouvoutsis, Constantinos Patsakis, and Fran Casino (University of Piraeus / associated research groups) published "Exploring the connection between coding habits and cognitive styles in malware developers" (arXiv:2606.05945v1, submitted 4 June 2026; covered by Help Net Security on 9 June 2026). The authors assembled a corpus of 658 leaked malware source-code projects from the VX-Underground repository and compared them against 249 benign open-source projects (including Python and JavaScript packages and security tooling such as nmap, sqlmap, and OWASP ZAP), analyzing them with four static application security testing (SAST) tools: Cppcheck for general C/C++ defects, Bandit for Python security issues, Snyk for dependency/package vulnerabilities, and Semgrep for pattern-based weaknesses.

The headline finding is that close to 90% of the analyzed malware projects contained at least one recognized software weakness. Poor code quality was the single most frequent category — missing integrity checks, unused/dead variables, and dead code dominated the results. Critically for defenders, a notable subset of samples shipped with TLS/SSL certificate validation disabled in their C2 client code (Improper Certificate Validation, CWE-295), which means an operator-in-the-path can intercept, decrypt, manipulate, or sinkhole the malware's command-and-control channel — an adversary-in-the-middle posture that the malware's own bug enables. More than 40% of the identified weaknesses involved code fragments shared across multiple malware families, indicating heavy code reuse and copy-paste among threat-actor codebases; the same exploitable bug therefore generalizes across families rather than being a one-off. Structural software metrics (cyclomatic complexity per function, maintainability index) were measured on the 463 of 658 projects (~70%) that supported structural measurement; malware showed smaller codebases, reduced documentation, higher per-function cyclomatic complexity, and minimal use of abstraction (classes, closures), consistent with development optimized for expedience, operational secrecy, and evasion rather than maintainability.

The study positions itself alongside the Malvuln project (cataloging exploitable bugs in malware since 2021). Operationally, the defensive takeaways are: (1) treat malware C2 clients with disabled certificate validation as interceptable — blue teams and takedown operators can perform TLS interception or impersonate C2 to enumerate, sinkhole, or disrupt; (2) shared/reused weaknesses mean a single detection or disruption technique can scale across families; and (3) the limitations matter — the corpus is C/C++-heavy and reflects only publicly leaked malware, so findings may not generalize to actively maintained, closed adversary toolchains. The named families (WannaCry, Emotet, Torpig) are illustrative of the reuse pattern and not the subject of new vulnerability disclosures here.

## MITRE ATT&CK

- T1071 Application Layer Protocol
- T1573 Encrypted Channel
- T1557 Adversary-in-the-Middle
- T1566 Phishing
- T1059 Command and Scripting Interpreter
- T1210 Exploitation of Remote Services
- T1056 Input Capture
- T1185 Browser Session Hijacking
- T1027 Obfuscated Files or Information
- T1497 Virtualization/Sandbox Evasion
- T1486 Data Encrypted for Impact
- T1490 Inhibit System Recovery
- T1568 Dynamic Resolution
- T1008 Fallback Channels
- T1095 Non-Application Layer Protocol
- T1105 Ingress Tool Transfer
- T1547 Boot or Logon Autostart Execution
- T1003 OS Credential Dumping
- T1497 Virtualization/Sandbox Evasion
- T1041 Exfiltration Over C2 Channel

## Sources

- [Malware ships with bugs that defenders could use against it](https://www.helpnetsecurity.com/2026/06/09/malware-source-code-bugs-research/)
- [Exploring the connection between coding habits and cognitive styles in malware developers (abstract)](https://arxiv.org/abs/2606.05945)
- [Exploring the connection between coding habits and cognitive styles in malware developers (PDF)](https://arxiv.org/pdf/2606.05945)
- [Malvuln — catalog of vulnerabilities in malware](https://malvuln.com/)
- [VX-Underground malware source-code repository (study dataset)](https://vx-underground.org/)
- [CWE-295: Improper Certificate Validation](https://cwe.mitre.org/data/definitions/295.html)
- [Analyzing SSL/TLS Certificates Used by Malware (corroborating C2 TLS research)](https://www.trendmicro.com/en_us/research/21/i/analyzing-ssl-tls-certificates-used-by-malware.html)

## Full data

Detection queries (Splunk SPL / Microsoft KQL / Sigma) and IOC values require the Threadlinqs MCP server (Purple tier): https://intel.threadlinqs.com/mcp

Canonical: https://intel.threadlinqs.com/threat/TL-2026-0739
