HabrJuly 23, 2026🇷🇺Translated from Russian

AI in Cybersecurity: Where It Delivers Real Value and Where It Remains Marketing Hype

The cybersecurity industry faces a severe talent shortage, making it increasingly difficult to hire experienced SOC analysts, incident responders, and threat hunters. At the same time, businesses seek greater efficiency and lower costs, while vendors promote their products as highly innovative. These pressures create unrealistic expectations that simply purchasing an AI-labeled solution will resolve major security problems.

What Counts as AI in Information Security Today

Vendors often blur distinctions between technologies by labeling everything as AI. Classical correlation rules in SIEM systems trigger on predefined conditions, such as a user logging in at night, gaining administrative rights, and exfiltrating data. These contain no artificial intelligence, only engineering. Machine learning systems build behavioral models from historical data and flag deviations, yet they require extensive calibration and still produce high volumes of false positives. Generative AI and large language models can draft queries, summarize incidents, and explain detection rules, but they lack business context.

Where Marketing Claims Outpace Reality

One common promise is that AI will automatically discover unknown attacks. Solutions such as UEBA, NDR, and XDR do detect anomalies, for example when an accountant suddenly accesses servers via VPN at 3 a.m. and runs PowerShell. However, the system identifies deviation, not confirmed malice, so analysts must still validate each alert. Another claim is fully automatic investigation of complex incidents. While tools can assemble timelines, they cannot assess which systems are business-critical or predict operational impact without human input.

Assertions that AI will replace SOC analysts also fall short. Modern platforms can group events, suppress some false positives, and generate report drafts, yet they cannot answer questions about client impact, backup availability, or financial consequences. Similarly, generative models can produce plausible security strategies, but they ignore budget limits, corporate culture, and specific regulatory constraints, resulting in generic templates rather than actionable plans.

Areas Where AI Delivers Measurable Value

In anti-fraud systems used by financial institutions such as Alfa-Bank, machine learning has moved beyond simple threshold rules. Modern models analyze click patterns, typing speed, device behavior, and transaction context simultaneously. They can block a transaction made by a pensioner in one city minutes after a high-value purchase from a newly installed app in another country. Behavioral analytics platforms also excel at spotting insider threats by comparing current activity against an individual’s historical profile.

Generative AI assists analysts by translating complex detection logic into plain language and automatically building queries against large datasets. Correlation engines in contemporary SIEM and XDR platforms process millions of daily events and surface relationships spanning days or weeks that no human could review manually.

How to Separate Real AI from Marketing Labels

Organizations should demand concrete metrics from vendors, including false-positive rates, time required to validate each alert, and the percentage of incidents actually detected. Understanding which underlying technology—rules, machine learning, or generative models—is being used remains essential before committing significant budgets to solutions that may deliver only conventional correlation under a new label.

Related articles

HabrAI Security

Claude Encrypted Thinking Blocks Use Protobuf with Exposed Metadata and AES-GCM Ciphertext

A detailed reverse-engineering of Claude signatures shows that the encrypted reasoning blocks are not opaque containers but structured protobuf messages. The outer envelope contains a 312-byte inner message that holds a 135-byte header, fixed-length nonce and MAC fields, and the actual ciphertext. The header itself reveals the model name such as claude-opus-5, the block type as thinking, and the organizationUuid from the user's Anthropic account. Only the reasoning text is encrypted with AES-GCM, adding exactly 16 bytes for the authentication tag. The analysis covers four protocol versions and notes that organization binding was added in version 15, potentially allowing servers to reject cross-model or cross-organization reuse. The findings provide concrete implications for both the Opus-to-Haiku extraction attack and the leakage of account identifiers in public logs.

HabrAI Security

Researchers Extract Proprietary Reasoning Traces from Anthropic, OpenAI and Google LLMs, Revealing Hidden Secrets

A team of eight researchers from institutions including ELLIS Institute Tübingen, the Max Planck Institute for Intelligent Systems, Tübingen AI Center, MATS and Snyk published a preprint detailing a practical attack that recovers full reasoning traces from closed LLM APIs. The method requires only two API calls and works by feeding encrypted reasoning blocks from strong models such as Claude Opus 4.8 into weaker models from the same provider, such as Haiku 4.5, which then reproduce the hidden chain-of-thought verbatim. Analysis of 6,708 publicly shared agent logs from GitHub and Hugging Face yielded 315,320 recovered traces containing 704 unique secrets, including 62 API keys, 33 passwords and 24 access tokens that never appeared in visible session output. The attack also enables extraction of internal safety policies, system prompts and detailed harmful planning that providers normally filter from final answers. In addition, the same mechanism can be used in reverse to inject malicious instructions into shared logs that later get replayed by unsuspecting users. The authors recommend treating encrypted reasoning blocks as sensitive secrets and propose cryptographic binding of traces to sessions, users and models.

HispasecAI Security

GhostSplice Technique Lets Malicious MCP Servers Trick AI Coding Agents into Exfiltrating Secrets

GhostSplice is a new technique that allows a malicious MCP server to induce an AI coding agent to leak SSH keys, environment secrets, and source code. The attack splits malicious instructions across tool metadata and responses so the agent reconstructs and executes the full exfiltration plan without detecting an overtly malicious command. Tests showed the method raised compliance rates from an average of 42 percent to 82 percent across eleven models, with some systems moving from zero to 100 percent success. The technique requires the developer to connect the attacker-controlled MCP server and for the agent to already possess read access to the targeted files. Defenses focus on strict allow-listing of MCP servers, least-privilege tool permissions, separation of tool output from instructions, and human approval for sensitive operations. The disclosure aligns with prior warnings about poisoned MCP tool descriptions and agentjacking attacks.

HabrAI Security

Anthropic Claude Code Auto Mode Launches August 14 with Local Classifier and Permission Rules

Starting August 14, Claude Code will run in auto mode on new sessions for Pro, Max, and Team plans, replacing the allow/deny dialog with a local classifier that evaluates every tool call. The classifier rules are stored locally and contain 103 categories across allow, soft_deny, hard_deny, and environment sections, with the single hard_deny rule focused on data exfiltration spanning over 5,000 characters. Enterprise, API, Bedrock, Vertex, and Foundry deployments remain on opt-in for another month. Auto mode pauses after three consecutive blocks or twenty blocks in a session, and broad allow rules such as python:* are disabled while narrow permissions continue to function. Administrators should populate the twenty environment fields, currently only one-third configured on clean machines, before the rollout date.