SentinelLabs Benchmark: Frontier LLMs Attempt Autonomous Long-Horizon Malware Analysis Using the 2005 Pre-Stuxnet 'fast16' Sabotage Toolkit as Case Study — Threadlinqs Intelligence
As of 2026-07-22, SentinelLabs Benchmark: Frontier LLMs Attempt Autonomous Long-Horizon Malware Analysis Using the 2005 Pre-Stuxnet 'fast16' Sabotage Toolkit as Case Study is a informational-severity threat intel threat attributed to Unknown (suspected nation-state, tracked by Threadlinqs Intelligence with 9 detection rules (Splunk SPL, Microsoft KQL, Sigma) and 20 indicators of compromise.
Threat ID: TL-2026-1630 · Severity: INFORMATIONAL · Status: ACTIVE · Category: THREAT_INTEL
Attribution: Unknown (suspected nation-state · ESPIONAGE
SentinelLabs ran an eight-stage benchmark testing whether frontier AI models (GPT-5.6 Sol, GPT-5.5, GLM-5.2, Claude Opus 4.6/4.7/4.8, Grok 4.2, Gemini 3.1 Pro, DeepSeek Reasoner) can autonomously
In April 2026, SentinelLabs researchers Vitaly Kamluk and Juan Andres Guerrero-Saade (with contributions from Silas Cutler and Costin Raiu) publicly disclosed 'fast16' at Black Hat Asia: a previously unknown Windows sabotage malware framework compiled in mid-to-late 2005, predating the public disclosure of Stuxnet by roughly five years. The discovery originated from a carrier executable, svcmgmt.exe, uploaded to VirusTotal in 2016 but bearing an August 2005 PE compile timestamp. svcmgmt.exe operates as a Windows service host that, depending on command-line arguments, can run as a service, execute an embedded Lua 5.0 interpreter (extended with a custom Unicode 'wstring' module and a built-in symmetric cipher routine for decrypting payload blocks), or interpret a filename argument to spawn two further commands. A companion boot-start kernel-mode driver, fast16.sys, implements a patch engine of 101 byte-pattern rules that locate specific code sequences inside targeted 32-bit Windows executables (Intel-compiler-linked .EXE files), capture original addresses, splice in replacement code -- including an injected floating-point-unit (FPU) block for 'precision arithmetic and scaling values in internal arrays' -- and repair PE metadata/checksums to avoid detection. This is not a rootkit in the conventional sense; SentinelLabs stresses the driver's function is highly targeted binary patching of computational logic, not file-hiding.
The likely targets identified via static/dynamic analysis and corpus testing are engineering/scientific simulation packages: LS-DYNA 970 (an explosion/impact/multiphysics finite-element solver used in structural and munitions modeling), PKPM (a Chinese structural-engineering CAD suite with the SATWE analysis engine), and MOHID (a Portuguese hydrodynamic/environmental modeling platform) -- the latter two sharing Fortran compiler lineage with LS-DYNA and triggering only 'haphazard' rule collisions rather than meaningful patches. The Institute for Science and International Security has documented Iran's historical use of LS-DYNA within its nuclear weapons-related engineering programs -- the same broader program later targeted overtly by Stuxnet -- which SentinelLabs cites as circumstantial evidence that fast16 may have been an early, covert attempt to sabotage nuclear-relevant physics simulations by introducing silent, systematic numerical errors rather than exfiltrating data or causing an overt crash. A reference to fast16-like tradecraft ('*** Nothing to see here - carry on ***') was found in the April 2017 Shadow Brokers 'Territorial Dispute' leak's drv_list.txt, a file believed to catalog NSA operator deconfliction/avoid-fratricide signatures, further supporting a nation-state origin. Source-control artifacts (SCCS/RCS style markers) embedded in the code point to developers working in long-term government- or military-grade Unix development environments.
Building on this real-world case, SentinelLabs designed an eight-stage benchmark to test whether current frontier LLMs could perform the equivalent long-horizon reverse-engineering investigation autonomously, end to end, without an operator doing the analytical work. Stages progressed from producing a gold-master IDA Pro IDB and identifying components (Stage 1), through patch-logic impact assessment on a single LS-DYNA installer (Stage 2), generalizing across a wider LS-DYNA corpus (Stage 3), correctly rejecting unsuitable candidates PKPM and MOHID (Stage 4), forming and testing hypotheses against unavailable ANSYS AUTODYN targets (Stage 5), validating claims against external public reporting (Stage 6), producing a publication-grade report (Stage 7), and reassessing conclusions under newly introduced ANSYS AUTODYN evidence (Stage 8). Models were graded on the quality of the underlying IDA Pro 9.3 databases they produced, not on prose summaries, and evaluated for local analytical soundness plus whether that quality held as investigative scope expanded ('inve
Target sectors: research, defense, energy, nuclear critical-infrastructure-engineering, artificial-intelligence
Target regions: iran, Global
Detections & IOCs
As of 2026-08-07, this threat has 9 detection rule(s) across Splunk SPL, Microsoft KQL and Sigma, and 20 indicator(s) of compromise. Detection query text and full IOC values are available to authenticated users and programmatically via the Threadlinqs MCP server (Purple tier). View plans.
THREAT_INTEL, INFORMATIONAL, threat intelligence, cybersecurity, T1543, T1547, T1036, T1027, T1140, T1562, T1070, T1518, T1083, T1082