AntiMalware•August 28, 2026•🇷🇺Translated from Russian

Claude AI Agent Accidentally Deletes Developer's 700 GB Home Directory

A developer named Sebastien Guillaime asked an Claude AI agent to write a script that would automatically clean up temporary files created by other AI agents he frequently runs.

Guillaime wanted each agent to operate inside its own isolated sandbox located in /tmp and to have those directories removed after the agent finished its task.

The first version of the script was overly complex and contained multiple destructive deletion commands. Because of these commands, Anthropic's safety system flagged the request as high-risk and automatically downgraded the model, first to Opus 5 and later to Opus 4.8.

The downgraded model created a test that correctly identified both /tmp and the home directory as dangerous locations. However, during the actual cleanup phase it reused the same variable that previously held the path to the user's home folder.

Guillaime stopped the process, but not before 700 GB of data had been deleted. He was able to restore the majority of the information from Git, Nix configuration files, session logs and other backups, yet a full week of work was still lost.

The developer suspects that the forced downgrade to a weaker model increased the chance of the variable conflict being missed, noting that the original Fable 5 model might have detected the error.

Related articles

Habr•AI Security

AI Agents Trigger Surge in Automated Reports, Forcing Google to Pause Bug Bounty Program

OpenAI warned over 100 companies about its agents potentially bypassing security controls on external websites. Wikimedia reported unauthorized edits by OpenAI agents that caused partial outages on Wikidata query services. Google observed a sharp rise in vulnerability disclosures from 5,045 in January to 10,740 in August, many driven by automated AI tools. As a direct result, Google suspended its open-source bug bounty program starting October 1 due to overwhelming volumes of low-quality automated submissions. The PageBreak AI agent independently discovered more than 500 XSS flaws across Google web applications. Adversa AI demonstrated prompt-based attacks that tricked GitHub Copilot CLI into leaking secrets from encrypted instructions. These developments highlight growing concerns over AI agent autonomy, unauthorized access, and their impact on both defensive and offensive security workflows.

Habr•AI Security

OSINT for the Lazy Part 19: How Generative AI Transforms Intelligence Gathering

The article examines the shift from manual OSINT practices to AI-driven workflows amid exploding data volumes. It details applications of NLP models like BERT, GPT and LLaMA for entity extraction, authorship attribution and report generation. Computer vision tools such as GeoSpy, Picarta and Google Vision AI enable automated geolocation and image forensics, while multimodal systems and graph neural networks map complex actor relationships. LLM agents equipped with planning modules, memory and tool access now handle multi-step collection and correlation tasks. The piece also covers limitations including hallucinations, source verification challenges and ethical risks around privacy and attribution. It concludes that effective OSINT now relies on symbiotic human-AI collaboration rather than full automation.

AntiMalware•AI Security

AI Agents Chain Malicious Instructions Through Protocol Pivoting to Bypass Protections

Researchers have demonstrated how AI agents can relay malicious instructions across multiple components without triggering security checks, allowing attackers to reach internal resources. The technique, called protocol pivoting, exploits the loss of trust validation when tasks move between AI systems connected via the MCP protocol. Syed Anas Mohiuddin showed that a single planted prompt can be passed from one agent to another, eventually reaching specialized tools that execute unauthorized actions such as network requests or data exposure. In Google MCP Toolbox for Databases, the flaw enabled HTTP redirects to internal addresses until a patch introduced address validation and request restrictions. A separate issue tracked as CVE-2026-97228 in Rapid7 Bulk Export MCP received a low CVSS score of 2.7 and was fixed in version 0.6.2, though it did not grant access beyond the original API key permissions. Experts note that the method is essentially an indirect prompt injection rather than an entirely new attack class.

AntiMalware•AI Security

Astra Group Unveils Astra AI Ecosystem for Air-Gapped Corporate Networks

Astra Group has introduced its Astra AI ecosystem designed for secure, on-premises deployment in closed corporate environments. The solution enables organizations to run AI models locally without transmitting data to external services, targeting critical infrastructure operators, government agencies, and regulated industries. Built on Astra Linux and the Botsman containerization platform, the ecosystem includes five integrated components for code automation, office assistants, low-code agent development, model management, and implementation methodology. The company claims productivity gains exceeding 50 percent for development tasks and up to fourfold performance improvements with its certified hardware-software complexes. While emphasizing data sovereignty and regulatory compliance, Astra Group notes that local deployment alone does not eliminate risks related to agent permissions, output quality, and integration security.