Ghostcommit Attack: Malicious Prompts Hidden in PNG Images Hijack AI Coding Agents to Steal .env Secrets
A new supply-chain attack technique named Ghostcommit enables adversaries to conceal prompt-injection commands inside ordinary PNG images, thereby bypassing AI code-review systems and compelling coding agents to exfiltrate sensitive secrets such as .env configuration files.
Attack Principle and Comparative Verification
The ASSET Research Group conducted experiments showing that when a pull request directly contains plaintext instructions to steal a repository’s .env file, tools such as the Cursor vulnerability bot and CodeRabbit immediately detect the malicious command and flag it as a high-severity issue before the code is merged.
To circumvent these safeguards, the researchers split the attack payload into two parts: a seemingly legitimate specification document named AGENTS.md placed in the repository root, and a PNG image (docs/images/build-spec.png) that actually contains the complete malicious logic. The image instructs the agent to read the .env file byte-by-byte and encode its contents as an ASCII numeric sequence. Because most text-oriented review tools treat images as opaque binary files and CodeRabbit skips PNG files by default, the entire pull request passes without any security warnings.
Attack Trigger Flow
The theft does not occur at merge time. Instead, the trap activates later when a developer, in a separate session, asks the coding agent to implement ordinary functionality. The agent consults the already-merged AGENTS.md file, parses the referenced image, reads the local .env file, and emits a module-level constant (for example _PROV_CANARY) containing hundreds of numeric values. When these numbers are decoded, they reconstruct the complete plaintext of the .env file. In one real-world test using Claude Sonnet inside Cursor, the model produced a tuple of 311 numbers that, once decoded, revealed API keys, database addresses, and cloud credentials.
Two Major Detection Blind Spots Exploited
- Malicious prompts are hidden inside images, rendering them invisible to human reviewers, regex scanners, and large-language-model tools that only analyze textual diffs.
- Extracted secrets are disguised as numeric tuples; conventional secret-scanning tools that rely on string patterns fail to recognize ASCII-decodable number sequences.
Multi-Tool and Multi-Model Test Results
Researchers evaluated 11 combinations of coding tools and models. Results showed that whether the attack succeeds depends primarily on the agent framework rather than the underlying model. Cursor leaked .env contents with Claude Sonnet 4.6, Composer-2, and GPT-5.5. Antigravity leaked secrets with Sonnet, Gemini 3.1 Pro, and Gemini 3 Flash. In contrast, the Claude Code platform blocked the attack across its entire model family (Sonnet 4.6, Haiku 4.5, Opus 4.7). Notably, the same Sonnet model behaved differently across environments. In one extreme test, Antigravity running Opus initially computed the secrets but then autonomously removed the sensitive output after recognizing the social-engineering pattern; the same model inside Claude Code never read the AGENTS.md file and therefore never triggered its self-defense mechanism.
Self-Developed Defense Prototype
The research team built a lightweight multimodal GitHub code-review tool based on the open-source Gemma 4 model that runs on a single 4 GB GPU. The prototype combines three detection capabilities: invisible-character identification, code-structure analysis, and simultaneous LLM-based validation of both specification documents and embedded images. In stress testing it achieved 100 % interception of 15 known attack samples with zero false positives, detected 49 out of 50 attacks in 80 previously unseen real-world submissions, and correctly classified all 30 normal business pull requests without raising alerts. The team has open-sourced the complete proof-of-concept code, including the split-payload pull-request samples and numeric-decoding utilities, to help the security community develop countermeasures.
Related articles
Zero Trust for AI Agents: Why Separate Identity Alone Is Not Enough
Denis Korbakov, CTO of Smart-Soft, explains why traditional IAM approaches fail to secure autonomous AI agents that dynamically select tools, change context, and delegate authority. Only 21.9% of teams treat agents as distinct identity-bearing entities, while 45.6% rely on shared API keys and 44.4% use generic tokens. Research from Gravitee, Cloud Security Alliance, and Aembit shows that 68% of organizations cannot distinguish AI agent actions from human actions, 74% grant excessive privileges, and 52% allow rights inheritance. The article maps NIST SP 800-207 Zero Trust principles—explicit verification, least privilege, and assume breach—to agent workloads using short-lived scoped tokens, SPIFFE/SPIRE credentials, and layered policy enforcement points. A concrete ticket-diagnosis scenario illustrates how prompt injection can be contained through per-task authorization, dedicated network segments, and independent telemetry from NGFW and SIEM. The piece concludes with an open question on sub-agent delegation chains and offers reference OPA/Rego policies plus runbooks for pilot implementations.
AWS Details Architecture to Reduce Prompt Injection Risks in AI Agents
AWS has introduced a new architecture designed to prevent compromised or manipulated AI agents from accessing data beyond user permissions. The approach relies on Amazon Bedrock AgentCore to shift authorization decisions from the agent itself to the underlying infrastructure and connected services. The core risk arises when agents receive broad credentials to query databases, repositories, and SaaS platforms, allowing potential prompt injection attacks to retrieve unauthorized information. In the proposed design, users authenticate via Amazon Cognito and receive JWT tokens containing attributes such as department or role. The AgentCore Runtime validates these tokens before executing any agent actions, rejecting requests that violate configured rules. For DynamoDB queries, temporary credentials are issued through AssumeRoleWithWebIdentity, with IAM policies enforcing strict access to authorized data partitions only.
Cybercriminals Weaponize OpenClaw AI Agent in ClawHavoc Campaign to Distribute Infostealers
Threat actors have repurposed the OpenClaw AI agent to deliver infostealers by uploading hundreds of malicious skills to ClawHub. The campaign, named ClawHavoc, tricks users into executing encoded commands or installing required tools under the guise of helpful AI recommendations. Researchers at Trellix identified 341 malicious skills, with 335 targeting installation of Atomic macOS Stealer on macOS systems. On Windows, victims receive password-protected archives and fake verification utilities that mirror classic ClickFix tactics. Analysis of repository history uncovered 1,184 suspicious packages linked to 12 authors, enabling theft of passwords, browser data, crypto wallets, API keys, SSH keys, and source code. Users are advised to update OpenClaw, audit installed skills, remove suspicious packages, and rotate potentially compromised credentials while running the agent in a restricted environment.
Server Log Analysis Reveals How Major AI Crawlers Actually Behave on Websites
A detailed examination of server access logs shows that AI vendors operate multiple distinct bots with separate purposes rather than a single crawler. GPTBot performs scheduled training data collection while OAI-SearchBot builds search indexes and ChatGPT-User fetches pages in direct response to user queries. The same pattern appears with PerplexityBot and Perplexity-User at Perplexity as well as ClaudeBot and user agents at Anthropic. Blocking all AI-related user agents in robots.txt therefore prevents both training crawls and live user-driven visits. Analysis of 515 million AI bot events found only 408 requests for llms.txt, confirming the file sees negligible adoption. Verification of IP addresses against vendor-published ranges remains the reliable method for distinguishing genuine bots from spoofed traffic. Effective practices focus on clean HTML structure, fast response times, and selective robots.txt rules that allow user-agent traffic while restricting training crawlers.