Researchers and security teams reported mixed results on using large language models for offensive cybersecurity tasks, finding that current systems remain unreliable for autonomous exploitation while still proving useful in vulnerability research. Google Project Zero's Project Naptime evaluated whether LLMs could meaningfully support offensive security work, while the CyberSecEval 2 benchmark measured model behavior across exploit generation, prompt injection, and code interpreter abuse. The benchmark found prompt-injection defenses remain weak, with successful attack rates of 26% to 41% across tested models, and highlighted a safety-utility tradeoff in which stronger refusals can also block benign requests.
At the same time, multi-agent LLM-assisted tooling demonstrated practical impact in software auditing. DARKNAVY said its Argusee architecture, which separates manager, auditor, and checker roles to reduce false positives and false negatives, uncovered 15 previously unknown flaws in open-source projects and identified CVE-2025-37891, a high-severity Linux kernel USB vulnerability. The flaw affects the host side of the Linux USB stack from Linux 6.5 onward, can be triggered by a simulated USB device advertising the MIDI2 protocol, and stems from improper length checks during MIDI1-to-UMP conversion that can cause an arbitrary kernel heap overflow; DARKNAVY reported reliable privilege escalation to root on Arch Linux before the issue was fixed.

Mallory correlates global threat intelligence with your attack surface — know if you’re exposed before adversaries strike.
9 events from the most recent confirmed update back to the earliest known activity.
At Black Hat in Las Vegas, James Kettle presented research showing that current agentic AI systems are limited at fully autonomous novel attack discovery but become highly effective with human guidance. He also disclosed a new web security attack surface he calls Shared-Parser Confusion, involving shared parsing logic for both untrusted requests and trusted responses.
DARKNAVY published a technical report describing DoGNAVY, an agentic vulnerability-reproduction system, and evaluated it on the full CyberGym Level 1 benchmark. The report said DoGNAVY passed differential validation on 1,369 of 1,507 tasks for a 90.84% success rate and included details on sandboxing, restricted network access, and a network-access audit.
ProjectDiscovery researcher Tarun Koyalwar published a behavioral audit of general-purpose AI models attacking a patched Argus validation benchmark with 54 usable black-box web targets. The study reported that benchmark solve rates masked execution failures, unintended exploit paths, and harness or side-channel interactions, and argued that behavioral auditing is needed to assess offensive-security capability.
The authors introduced CyberGym, a large-scale benchmark for evaluating AI agents on realistic cybersecurity tasks across 1,507 real-world vulnerabilities in 188 software projects. The paper reported that top-performing agent-model combinations achieved only about a 20% success rate and said use of CyberGym led to the discovery of 34 zero-day vulnerabilities and 18 historically incomplete patches.
DARKNAVY stated it exploited CVE-2025-37891 to reliably escalate privileges to root on Arch Linux. The same disclosure noted that the vulnerability had been fixed and referenced a Linux kernel stable patch.
DARKNAVY reported that Argusee discovered CVE-2025-37891, a high-severity Linux kernel USB subsystem vulnerability affecting distributions including Ubuntu and Arch Linux. The flaw is triggered by a simulated USB MIDI2 device and stems from improper length checks that can lead to an arbitrary kernel heap overflow.
DARKNAVY described Argusee, a multi-agent LLM-assisted vulnerability discovery system, and reported that it found 15 previously unknown flaws in medium-sized open-source projects including GPAC and GIFLIB. The post also said Argusee achieved strong benchmark results, including 100% accuracy on some CyberSecEval 2 categories.
On submission, the authors presented CyberSecEval 2 as a benchmark for measuring large language model security risks and cybersecurity capabilities, adding evaluation areas for prompt injection and code interpreter abuse. The paper also reported prompt-injection success rates of 26% to 41% across tested models and released the evaluation code as open source.
1Password's Off-By-1-Labs published research evaluating frontier AI models on patching six recently disclosed open-source vulnerabilities, finding that only 26% of 6,080 generated patches were usable while 53.9% failed to fix the issue, introduced new bugs, or both. The researchers characterized many outputs as 'FLAWED' patches and released the FLAWED tooling on GitHub for further study.
Vulnerabilities, threat actors, malware, products, organizations, and breaches Mallory has linked to this story.
8 references tracked. Mallory keeps watching after this page renders.
zdnet.com
Open sourcewired.com
Open sourcedarknavy.org
Open sourceprojectdiscovery.io
Open sourcearxiv.org
Open sourcedarknavy.org
Open sourcegoogleprojectzero.blogspot.com
Open sourcearxiv.org
Open sourceMap indicators from this story to your assets and identify affected systems in minutes.
Every observed campaign, victim, and pivot linked to actors named in this story.
Malware, exploits, and IOCs connected to the activity described here.
YARA, Sigma, and Snort rules deployed to your SIEM as soon as they’re published.
Get matching new stories delivered to your team as they break — not the next morning.
Ask questions about this story and take action on the answers.