Automated Pentesting and BAS: How AI Systems Like XBOW Outpace Human Researchers in Vulnerability Discovery
The rapid evolution of infrastructure demands continuous security validation rather than annual manual pentests. Automated pentesting and Breach and Attack Simulation (BAS) tools now enable organizations to test defenses daily by safely replaying attacker behaviors.
BAS platforms such as Cymulate, SafeBreach, AttackIQ, and Picus focus on individual techniques mapped to the MITRE ATT&CK matrix. They verify whether antivirus, SIEM, or EDR solutions detect and block actions like privilege escalation or lateral movement without searching for new vulnerabilities.
Autopentest solutions go further by chaining multiple steps into realistic attack paths that start from an entry point and reach critical assets. This approach answers whether an attacker can achieve an unacceptable event, providing direct feedback into vulnerability management workflows.
In Russia, Positive Technologies launched PT Dephaze with its commercial version released in February 2025. Version 3.0, released on 31 October 2025, added expanded attack chains, detailed evidence collection, and machine learning for fully controlled internal testing. The product received the National Runet Award in the Information Security category in December 2025. Other local offerings include APT BEZDNA from CTRLHACK and continuous pentesting platforms from BI.ZONE.
Globally, mature platforms such as Pentera and Horizon3.ai NodeZero already serve large enterprises. The most striking development occurred in 2025 when AI systems surpassed human researchers. Autonomous system XBOW accumulated over 2,000 reputation points and submitted more than 1,000 vulnerability reports, including 54 critical findings, securing first place on HackerOne within 90 days. In one benchmark, XBOW completed tasks in 28 minutes that took a 20-year veteran pentester 40 hours.
Google’s Big Sleep project, a collaboration between DeepMind and Project Zero, discovered CVE-2025-6965, a critical memory vulnerability in SQLite affecting all versions prior to 3.50.2. The team predicted imminent exploitation and coordinated a patch release within 48 hours, marking the first documented case of an AI agent preventing real-world exploitation.
These capabilities are being consolidated under Gartner’s new Automated Exposure Validation category, which merges BAS, autopentest, and external attack surface management. The findings loop directly back into vulnerability management, raising the priority of previously missed attack paths and ineffective controls.
Related articles
Critical SSRF Vulnerability Affects SonicWall SMA1000 Series Remote Access Appliances
SonicWall has disclosed four vulnerabilities in its SMA1000 series remote access products, with one rated critical. The most severe issue, CVE-2026-102255, is a server-side request forgery flaw in the WorkPlace interface that allows unauthenticated attackers to abuse the appliance as a forward proxy and reach internal functions. The vulnerability received the maximum CVSSv3.0 base score of 10.0. Two additional flaws, CVE-2026-102256 and CVE-2026-102257, enable authenticated OS command injection and unauthenticated path traversal via crafted archives, respectively. No exploitation has been observed in the wild at the time of disclosure. SonicWall has released updates to address all issues.
WordPress 7.1.3 Security Release Fixes Seven Vulnerabilities Including Stored XSS and SQL Injection
The WordPress development team has released version 7.1.3 as a maintenance and security update addressing multiple vulnerabilities. The release includes seven security fixes and four additional bug corrections. Among the security issues resolved is a stored cross-site scripting flaw that allowed pending comments to execute scripts in the administrative interface. Other fixes cover a denial-of-service condition in URL handling, an SQL injection vulnerability in the WXR export feature, and unauthorized disclosure of comments attached to private or unpublished posts. Additional patches address an XSS issue in the Imgur embed functionality, improper sticky post permissions for users with the Author role, and a parameter manipulation problem affecting hook action names.
Google Releases Chrome 155 Fixing 247 Vulnerabilities Including Four Critical Use-After-Free Flaws
Google has released Chrome 155 on October 6, 2026, addressing a total of 247 security issues across Windows, macOS, and Linux platforms. The update includes four critical vulnerabilities, all classified as use-after-free flaws that affect Chromecast, Browser, Navigation, and Track components. Fifty-three high-severity issues were also resolved, covering problems in SiteIsolation, Core, Omnibox, FileSystem, ANGLE, WebGL, and multiple other modules. The critical CVEs fixed are CVE-2026-106382, CVE-2026-106197, CVE-2026-106358, and CVE-2026-106347. Additional fixes address use-after-free conditions, race conditions, type confusion, and integer overflows in V8, WebRTC, PDF, Media, Parser, Storage, and WebAudio. The new versions are Chrome 155.0.8059.40 for Windows and macOS and 155.0.8059.39 for Linux and macOS.
LibreOffice and Apache OpenOffice Flaws Enable Remote Code Execution via Malicious Spreadsheets Without Macro Warnings
Two vulnerabilities, CVE-2026-63277 in LibreOffice Calc and CVE-2026-59265 in Apache OpenOffice, allow attackers to execute arbitrary code simply by tricking users into opening specially crafted spreadsheet files. The flaws exploit Java integration and class path handling, bypassing traditional macro security prompts entirely. LibreOffice has already released fixes in versions 26.2.5 and 26.8.0 that restrict class path entries to local file URLs only. Apache OpenOffice 4.1.16 and earlier remain vulnerable, with the stable patch expected in 4.1.17; interim mitigation requires disabling Java integration. The issues highlight risks in office suites that process untrusted documents containing external data connections or JDBC references. Organizations are advised to enforce least-privilege execution and avoid opening files from unknown sources until patches are applied.