HabrAugust 14, 2026🇷🇺Translated from Russian

Anthropic Rolls Out Invisible Statistical Watermarks for Claude Models to Comply with EU AI Act

On 2 August 2026, the same day Article 50 of the EU AI Act entered into force, Anthropic began embedding an invisible statistical watermark in every text generated by its latest Claude models, including Opus 5 and Sonnet 5. The company presented the measure as compliance with the voluntary Code of Practice on AI content transparency endorsed by the European Commission.

The watermarking system is two-layered. For text, a secret key known only to Anthropic partitions the model’s candidate tokens at each step into “green” and “red” groups and slightly increases the probability of selecting from the green group. Over hundreds of tokens this creates a detectable statistical bias that cannot be reproduced by chance without knowledge of the key. For image and document files the system uses signed C2PA metadata, the same open standard already used in photojournalism.

Although the regulation applies only to content served to European users, Anthropic chose to activate the watermark globally. Maintaining separate model versions for different jurisdictions would have imposed significant engineering overhead, so a single worldwide deployment was implemented.

Within 24 hours of the announcement, GitHub projects such as watermarks-remover by Guillaume Meyer appeared, quickly attracting thousands of stars. Commercial services including StealthGPT, claudewatermark.com and Human Writes also added “Claude watermark removal” features. Independent analysis of the released code showed that these tools currently address only Unicode artifacts and file metadata; none have demonstrated reliable removal of the statistical token bias.

The underlying technique is not new. It was introduced in the 2023 paper by Kirchenbauer et al. and later productised by Google as SynthID. OpenAI researcher Scott Aaronson described a similar approach in 2023, although it has not been deployed in ChatGPT. Because Anthropic has not published its detector or key-rotation policy, independent verification of removal claims remains impossible.

Research published at ICLR 2024 and NeurIPS 2024 confirms that even light paraphrasing or translation destroys most of the signal, while heavy rewriting reduces detection to random chance. Conversely, very short or rigidly structured outputs such as code snippets or tables contain too few token-choice opportunities for a reliable watermark to form.

Related articles

HabrAI Security

How to Interact with AI Models Without Exposing Sensitive Data

The article provides practical guidance on minimizing data leakage risks when using popular AI chatbots such as ChatGPT, Gemini, Claude and GigaChat. It explains that conversations are routinely scanned by automated filters and may be reviewed by human moderators or shared with law enforcement upon request. Key recommendations include disabling model training on user data, replacing sensitive values with placeholders, regularly deleting chat histories and verifying downloaded models for malicious injections. The guide also demonstrates local deployment using Ollama and secure API integration through the ChatBox client with Cloud.ru’s Evolution Foundation Models service. Local execution in Docker containers is presented as the most private option, although it requires significant computational resources. The author stresses that even after disabling training, data may still reach moderators and that users remain responsible for their own information.

HabrAI Security

Raft Develops Multilabel Guardrail Classifier Detecting 15 Risk Categories with 3x Speed and Cost Gains

Raft has released details on a custom multilabel guardrail classifier designed to scan both incoming prompts and model outputs for 15 distinct risk categories in real time. The system handles Russian and English text while maintaining independent thresholds for each category to balance false positives against critical misses. By switching from a PyTorch baseline to TensorRT inference on NVIDIA RTX 3090 hardware, the team achieved a 2.95x reduction in single-request latency and lowered inference cost to $0.062 per million requests. The architecture uses per-category expert query tokens plus a lightweight interaction transformer to capture correlations such as those between armament and violent content. Training relied on asymmetric loss functions and post-epoch per-class threshold tuning rather than standard binary cross-entropy. Benchmarking against nine open guardrail and toxicity models showed superior macro-F1 on rare but high-impact categories while remaining an order of magnitude cheaper than external LLM judges.

安全客AI Security

Zhou Hongyi Warns AI Tools Are Industrializing Vulnerability Discovery

At the Fourth Cyberspace Security Forum in Tianjin, 360 founder Zhou Hongyi stated that vulnerability mining is shifting from artisanal workshops to automated production lines, compressing discovery cycles from months or years down to hours. AI tools such as Mythos are standardizing and automating the process, enabling attackers to replicate elite hacker expertise at scale through distilled models and agent swarms. 360's own Tulongfeng platform has already discovered over 10,000 vulnerabilities since its June release, including long-hidden high-risk flaws in Windows, Office, OpenClaw, Flowise, and Codex. The emergence of multi-agent systems introduces new attack surfaces because compromised agents can autonomously collaborate and move laterally faster than human operators. Zhou described this as the "second one-way transparency," where offensive tradecraft becomes copy-pasteable via prompts and toolchains. Defenders are advised to adopt "model-versus-model" strategies, automate vulnerability intelligence workflows with SOAR, enforce strict agent permission audits, and integrate AI into their own code review and detection engineering processes.

安全客AI Security

Anthropic Fable 5.1 System Prompt Fully Leaked Hours After Launch Exposing 275000 Characters of Rules

Anthropic released its flagship Fable 5.1 model alongside Mythos 5.1 on September 2, achieving strong benchmark scores including 90 percent on ARC-AGI-2. Within hours, researcher Pliny the Liberator published the complete 275000-character system prompt on GitHub, far exceeding the company's official 27000-word disclosure. The leaked document details 46 built-in tools, strict copyright restrictions, memory classification boundaries, and behavioral constraints that function as an internal employee handbook. The incident highlights that model weights remain the true core while prompt-based guardrails create an attack surface once mapped. It also reveals privacy rules that permanently exclude storage of minor identities, criminal records, and self-harm indicators even when users disclose them. The leak underscores the growing gap between vendor transparency claims and actual runtime instructions governing frontier AI systems.