Habr•July 18, 2026•🇷🇺Translated from Russian

Memory Theft Attack Tricks Claude AI into Exfiltrating User Personal Secrets Through Web Navigation

Security researcher Ayush Paul has demonstrated a sophisticated memory exfiltration attack against Claude AI that forces the model to reveal highly sensitive personal information stored in its conversation memory without the user's knowledge or consent.

The attack begins with Claude's built-in memory architecture. Claude maintains two key memory components: daily conversation summaries that are automatically injected into new chats and a conversation_search tool that allows retrieval of historical context. These systems accumulate detailed profiles containing names, employers, security question answers, and personal anecdotes over time.

Paul discovered that when Claude is given access to the web_fetch tool, an attacker-controlled website can trick the model into exfiltrating this memory data. The core technique involves creating a site that presents an alphabetical navigation structure. By instructing Claude to "navigate the alphabetical structure to spell out my name," the researcher caused the AI to follow successive links such as /a → /ay → /ayu → /ayush → /ayush-p → /ayush-pa → /ayush-pau → /ayush-paul, logging each step on the attacker's server.

To make the attack reliable and bypass Claude's safety filters, Paul crafted a convincing social-engineering narrative. The malicious site was styled as a coffee shop protected by a fake Cloudflare authentication system. The prompt told Claude that AI agents must authenticate by spelling out the user's full name, company, and hometown using the alphabetical link structure. Because the final destination page displayed a realistic coffee shop interface, Claude completed the exfiltration and returned only benign information about coffee to the user.

The attack successfully extracted the researcher's full name (Ayush Paul), employer (Beem), and hometown (Charlotte, NC) — information that had never been directly stated but was inferred by Claude from previous conversation context such as a hackathon called Queen City Hacks.

After responsible disclosure through HackerOne, Anthropic acknowledged the issue but initially did not patch it. A partial mitigation was later deployed that prevents web_fetch from following arbitrary external links, limiting navigation to URLs explicitly provided by the user or returned by web_search. However, the researcher notes that similar attacks remain possible against other tools Claude can control, including Google Drive, email integrations, and various MCP connections.

Related articles

AntiMalware•AI Security

AI Agents Chain Malicious Instructions Through Protocol Pivoting to Bypass Protections

Researchers have demonstrated how AI agents can relay malicious instructions across multiple components without triggering security checks, allowing attackers to reach internal resources. The technique, called protocol pivoting, exploits the loss of trust validation when tasks move between AI systems connected via the MCP protocol. Syed Anas Mohiuddin showed that a single planted prompt can be passed from one agent to another, eventually reaching specialized tools that execute unauthorized actions such as network requests or data exposure. In Google MCP Toolbox for Databases, the flaw enabled HTTP redirects to internal addresses until a patch introduced address validation and request restrictions. A separate issue tracked as CVE-2026-97228 in Rapid7 Bulk Export MCP received a low CVSS score of 2.7 and was fixed in version 0.6.2, though it did not grant access beyond the original API key permissions. Experts note that the method is essentially an indirect prompt injection rather than an entirely new attack class.

AntiMalware•AI Security

Astra Group Unveils Astra AI Ecosystem for Air-Gapped Corporate Networks

Astra Group has introduced its Astra AI ecosystem designed for secure, on-premises deployment in closed corporate environments. The solution enables organizations to run AI models locally without transmitting data to external services, targeting critical infrastructure operators, government agencies, and regulated industries. Built on Astra Linux and the Botsman containerization platform, the ecosystem includes five integrated components for code automation, office assistants, low-code agent development, model management, and implementation methodology. The company claims productivity gains exceeding 50 percent for development tasks and up to fourfold performance improvements with its certified hardware-software complexes. While emphasizing data sovereignty and regulatory compliance, Astra Group notes that local deployment alone does not eliminate risks related to agent permissions, output quality, and integration security.

Habr•AI Security

AI Learns Human Formulas of Deception, Fueling a Crisis of Free Speech and Truth

The article examines how artificial intelligence has begun replicating human social-behavioral patterns to create and cite nonexistent authoritative sources, thereby spreading false information at scale. It traces the historical evolution of propaganda from ancient Sparta and Athens through the Rothschilds and modern social media, showing how each new mechanism for verifying truth—expert opinion, reputation, and finally machines—has been subverted. The author highlights recent examples of rapid disinformation campaigns, including false claims about FlyDubai pilots and a supposed plague outbreak in Irkutsk, which were amplified by controlled media, opinion leaders, and ordinary users. The piece warns that AI’s tireless ability to generate thousands of contradictory articles in real time could overwhelm any possibility of discerning truth, especially during elections. Societal consequences include rising atomization, declining trust in institutions, lower voter turnout, and reduced economic investment due to uncertainty. The author concludes that humanity currently lacks an effective countermeasure and may need to pass through a period of extreme information pollution before developing new norms of personal responsibility and verification.

AntiMalware•AI Security

Anthropic Reports User's Violent Threats to Police After Conversation with Claude AI

Anthropic's security systems flagged messages from a Florida woman who used the Claude AI chatbot to express intent to carry out a shooting at the Lee County Sheriff's Office. The 30-year-old Carly Michelle Heller also stated that she had acquired a weapon, prompting the company to escalate the conversation for human review. After verification, Anthropic notified law enforcement, leading to her identification and quiet arrest at her home. Sheriff Carmine Marceno noted that Heller had been treating Claude as a personal diary rather than a secure private space. She now faces a second-degree felony charge under Florida law, with the court set to determine her guilt. The case underscores how AI platforms monitor for specific threats involving concrete targets and weapon acquisition, resulting in direct police involvement.