Hermes Emerges as Modular Harness for Practical AI Security Testing
The language model by itself possesses no inherent security capabilities. The earlier this illusion disappears, the sooner a functional tool can appear. Running any frontier model out of the box and asking it to test an agentic system or craft a current attack vector produces instantly disappointing results. The model outputs naïve jailbreak attempts reminiscent of 2023, long after the industry moved past GPT-3.5-era thinking.
In real AI security work, base models remain anchored to stale weights, confuse hallucinations with actual vulnerabilities, and know nothing about recent long-term memory poisoning techniques, sandbox escapes, or fresh exploits in model repositories. The author, who runs the PWN AI Telegram channel, stresses that the field changes weekly. Without external scaffolding, the model stays a disconnected text generator. Only the combination of a language model and a specialised harness produces a dependable, reproducible agent.
After evaluating monolithic frameworks such as OpenClaw, the author rejected them because of excessive abstractions, token waste, and state leakage. The choice fell on Hermes for its modularity, dynamic skill loading, and minimal overhead. The harness runs on DeepSeek-v4-flash (April and July releases) chosen for strong agentic performance and low cost. A separate DeepSeek harness was tested but found too immature for production use.
The architecture rests on three control loops plus an execution layer. The management layer pins the agent to version-controlled runbooks, skills and plans stored in Git so every step remains auditable and reversible. A background data-collection layer continuously queries more than thirty sources—NVD, CISA KEV, Exploit-DB, Unit 42, Resecurity, arXiv and others—writing timestamped artefacts to disk. An analysis layer applies Hermes skills that clean noise, rank threats with the custom TIPS metric, and produce compact JSON summaries.
Execution occurs inside a restricted Docker sandbox connected to security tools including Garak, PyRIT, promptfoo, fickling, modelscan and presidio. Context inflation is fought through dynamic skill injection, report deduplication, and replacement of raw logs with structured JSON. Eight mandatory validation gates must all turn green before the model is allowed to emit a final answer.
Academic papers undergo a four-stage pipeline: candidate discovery via arXiv and Semantic Scholar, expert filtering, deep technical extraction by the Paper-to-PoC skill, and vision analysis of diagrams. Surviving findings are reproduced against a local test model under strict network and command restrictions. Successful reproductions are automatically committed back to the repository, allowing the harness to learn from its own results.
Related articles
Do Sandbox Restrictions Actually Work for AI Agents Running in Linux and gVisor?
An in-depth technical analysis examines whether security mechanisms such as Landlock, classic BPF socket filters, and CGROUP_DEVICE programs enforce intended restrictions inside container and VM-based sandboxes used by AI agents. Tests conducted on Linux 6.8 and two gVisor releases (20260817.0 and 20260831.0) revealed that Landlock calls consistently return ENOSYS inside gVisor, rendering the mechanism unavailable. CGROUP_DEVICE programs could be loaded and attached successfully under elevated capabilities, yet they produced no observable effect on device access. Classic BPF filters attached via SO_ATTACH_FILTER were accepted without error even with zero capabilities, but continued to allow UDP datagrams that should have been dropped. The study emphasizes that successful configuration alone does not guarantee enforcement and outlines a verification workflow that must be repeated for each target environment, runtime, and policy change before deploying restricted AI tools.
Houlong Security Industry Research Institute Releases 2026 China Cybersecurity Industry Map
The Houlong Security Industry Research Institute has published its comprehensive 2026 Network Security Industry Map following months of research that collected over 400 valid responses from leading Chinese cybersecurity firms. The report documents a structural market shift driven by AI-enabled attacks moving from theory to real-world operations, including automated phishing, deepfake fraud, and dual ransomware-extortion models targeting APIs and supply chains. On the defense side, it highlights the rapid adoption of AI for real-time threat detection, large-scale zero-trust deployments, privacy-preserving computation, and preparations for quantum-safe migration. The study notes that vendors integrating AI capabilities are outperforming peers in customer retention and pricing power while the industry moves away from broad product suites toward specialized, scenario-focused solutions. Overall, the map identifies three irreversible trends: AI becoming mandatory in security products, competition favoring depth over breadth, and sustained growth fueled by digital transformation and geopolitical factors.
Natalia Kaspersky Questions Trustworthiness Criteria for Generative AI
Natalia Kaspersky has expressed serious doubts about applying traditional trust criteria to generative AI systems. She explained that a trusted system must operate within predefined parameters and deliver predictable, repeatable results. Generative AI fails this standard because it produces varying outputs for the same inputs. The enormous scale of modern models makes comprehensive verification practically impossible. Selective testing of individual responses provides no assurance of overall reliability. Kaspersky stressed that creating trusted AI requires joint efforts from AI specialists, information security experts, methodologists, and standards developers rather than discussions alone.
Why 'You Are My Grandmother' Jailbreaks Succeed Against LLMs and How an External Controller Could Fix Them
The article examines why simple role-playing prompts easily bypass safety rules in large language models. It contrasts two possibilities: models that merely reproduce refusal templates versus those that maintain a stable internal representation of prohibited categories. Because competing contextual signals often outweigh safety constraints, jailbreaks succeed by shifting token prediction priorities. The proposed remedy separates the main LLM from an independent controller module that inspects both full input context and generated output against a narrow list of disallowed topics such as fraud, weapons, and child exploitation material. Several efficiency techniques are suggested, including block-wise scanning, embedding-based pre-filters, and two-stage checks that avoid reprocessing entire 100k-token dialogues on every turn. The author stresses that the controller must remain an external, non-LLM component to prevent recursive oversight layers. The discussion concludes that only such architectural separation offers robust resistance to context-based jailbreaks.