HabrAugust 23, 2026🇷🇺Translated from Russian

Hermes Emerges as Modular Harness for Practical AI Security Testing

The language model by itself possesses no inherent security capabilities. The earlier this illusion disappears, the sooner a functional tool can appear. Running any frontier model out of the box and asking it to test an agentic system or craft a current attack vector produces instantly disappointing results. The model outputs naïve jailbreak attempts reminiscent of 2023, long after the industry moved past GPT-3.5-era thinking.

In real AI security work, base models remain anchored to stale weights, confuse hallucinations with actual vulnerabilities, and know nothing about recent long-term memory poisoning techniques, sandbox escapes, or fresh exploits in model repositories. The author, who runs the PWN AI Telegram channel, stresses that the field changes weekly. Without external scaffolding, the model stays a disconnected text generator. Only the combination of a language model and a specialised harness produces a dependable, reproducible agent.

After evaluating monolithic frameworks such as OpenClaw, the author rejected them because of excessive abstractions, token waste, and state leakage. The choice fell on Hermes for its modularity, dynamic skill loading, and minimal overhead. The harness runs on DeepSeek-v4-flash (April and July releases) chosen for strong agentic performance and low cost. A separate DeepSeek harness was tested but found too immature for production use.

The architecture rests on three control loops plus an execution layer. The management layer pins the agent to version-controlled runbooks, skills and plans stored in Git so every step remains auditable and reversible. A background data-collection layer continuously queries more than thirty sources—NVD, CISA KEV, Exploit-DB, Unit 42, Resecurity, arXiv and others—writing timestamped artefacts to disk. An analysis layer applies Hermes skills that clean noise, rank threats with the custom TIPS metric, and produce compact JSON summaries.

Execution occurs inside a restricted Docker sandbox connected to security tools including Garak, PyRIT, promptfoo, fickling, modelscan and presidio. Context inflation is fought through dynamic skill injection, report deduplication, and replacement of raw logs with structured JSON. Eight mandatory validation gates must all turn green before the model is allowed to emit a final answer.

Academic papers undergo a four-stage pipeline: candidate discovery via arXiv and Semantic Scholar, expert filtering, deep technical extraction by the Paper-to-PoC skill, and vision analysis of diagrams. Surviving findings are reproduced against a local test model under strict network and command restrictions. Successful reproductions are automatically committed back to the repository, allowing the harness to learn from its own results.

Related articles

HabrAI Security

Zero Trust for AI Agents: Why Separate Identity Alone Is Not Enough

Denis Korbakov, CTO of Smart-Soft, explains why traditional IAM approaches fail to secure autonomous AI agents that dynamically select tools, change context, and delegate authority. Only 21.9% of teams treat agents as distinct identity-bearing entities, while 45.6% rely on shared API keys and 44.4% use generic tokens. Research from Gravitee, Cloud Security Alliance, and Aembit shows that 68% of organizations cannot distinguish AI agent actions from human actions, 74% grant excessive privileges, and 52% allow rights inheritance. The article maps NIST SP 800-207 Zero Trust principles—explicit verification, least privilege, and assume breach—to agent workloads using short-lived scoped tokens, SPIFFE/SPIRE credentials, and layered policy enforcement points. A concrete ticket-diagnosis scenario illustrates how prompt injection can be contained through per-task authorization, dedicated network segments, and independent telemetry from NGFW and SIEM. The piece concludes with an open question on sub-agent delegation chains and offers reference OPA/Rego policies plus runbooks for pilot implementations.

BoletimSecAI Security

AWS Details Architecture to Reduce Prompt Injection Risks in AI Agents

AWS has introduced a new architecture designed to prevent compromised or manipulated AI agents from accessing data beyond user permissions. The approach relies on Amazon Bedrock AgentCore to shift authorization decisions from the agent itself to the underlying infrastructure and connected services. The core risk arises when agents receive broad credentials to query databases, repositories, and SaaS platforms, allowing potential prompt injection attacks to retrieve unauthorized information. In the proposed design, users authenticate via Amazon Cognito and receive JWT tokens containing attributes such as department or role. The AgentCore Runtime validates these tokens before executing any agent actions, rejecting requests that violate configured rules. For DynamoDB queries, temporary credentials are issued through AssumeRoleWithWebIdentity, with IAM policies enforcing strict access to authorized data partitions only.

AntiMalwareAI Security

Cybercriminals Weaponize OpenClaw AI Agent in ClawHavoc Campaign to Distribute Infostealers

Threat actors have repurposed the OpenClaw AI agent to deliver infostealers by uploading hundreds of malicious skills to ClawHub. The campaign, named ClawHavoc, tricks users into executing encoded commands or installing required tools under the guise of helpful AI recommendations. Researchers at Trellix identified 341 malicious skills, with 335 targeting installation of Atomic macOS Stealer on macOS systems. On Windows, victims receive password-protected archives and fake verification utilities that mirror classic ClickFix tactics. Analysis of repository history uncovered 1,184 suspicious packages linked to 12 authors, enabling theft of passwords, browser data, crypto wallets, API keys, SSH keys, and source code. Users are advised to update OpenClaw, audit installed skills, remove suspicious packages, and rotate potentially compromised credentials while running the agent in a restricted environment.

HabrAI Security

Server Log Analysis Reveals How Major AI Crawlers Actually Behave on Websites

A detailed examination of server access logs shows that AI vendors operate multiple distinct bots with separate purposes rather than a single crawler. GPTBot performs scheduled training data collection while OAI-SearchBot builds search indexes and ChatGPT-User fetches pages in direct response to user queries. The same pattern appears with PerplexityBot and Perplexity-User at Perplexity as well as ClaudeBot and user agents at Anthropic. Blocking all AI-related user agents in robots.txt therefore prevents both training crawls and live user-driven visits. Analysis of 515 million AI bot events found only 408 requests for llms.txt, confirming the file sees negligible adoption. Verification of IP addresses against vendor-published ranges remains the reliable method for distinguishing genuine bots from spoofed traffic. Effective practices focus on clean HTML structure, fast response times, and selective robots.txt rules that allow user-agent traffic while restricting training crawlers.