Do You Really Know What Your AI Agent Is Doing in the Sandbox?
The rapid adoption of AI agents equipped with shells, filesystems, and network access has created a fundamental observability problem. When an agent finishes a task, logs generated by the model itself and platform metrics provide only a partial picture. The real question remains: what system calls did the agent actually execute?
Two main approaches exist for observing agent behavior. The first intercepts tool calls inside the agent framework, which is inexpensive but blind to any actions performed outside the framework. The second uses eBPF probes on the host kernel to capture every syscall, an approach used by container security tools. This method is more trustworthy yet assumes the agent shares the host kernel.
Strong isolation runtimes break that assumption. gVisor implements its own kernel in user space, while Kata Containers and Firecracker run a full guest Linux kernel inside a virtual machine. In both cases, guest syscalls never reach the host kernel where traditional monitors attach.
Existing standards cover several aspects of agent operation. MCP standardizes how agents connect to tools and data. OpenTelemetry defines semantic conventions for model and agent spans. RuntimeClass lets Kubernetes switch between runtimes without changing application manifests. The new agent-sandbox project treats sandboxes as first-class Kubernetes resources. None of these standards, however, define a common interface for reporting the actual syscalls performed inside the sandbox.
Empirical testing used a generator that performed ten distinct operations, each tagged with a unique marker. Probes verified that the monitoring stack itself was functioning correctly. Results across 24 runs demonstrated that visibility is determined by whether filesystem operations cross into a host process. When a block device backs the guest root, almost no guest file operations appear on the host. When virtiofs or an overlay shares the filesystem via a host process, file names and writes become visible.
Even within the same runtime, configuration changes produce dramatically different results. Enabling the in-memory overlay in one gVisor mode hides writes and deletions while still exposing opens. Executions never surface as execve events on the host under any tested sandbox; only prior file copies leave detectable traces when the filesystem is host-mediated.
Commercial security products already reflect these limitations. Falco removed its gVisor engine, Google Container Threat Detection is incompatible with GKE Sandbox, and Microsoft Defender cannot inspect Kata pods. The measurements confirm that switching a RuntimeClass value can silently eliminate visibility into entire classes of agent actions.
Related articles
AI Reshapes Cybersecurity Jobs: Automation of Routine Tasks, Rising Demand for Architects and AI Defenders
The cognitive revolution driven by AI technologies is transforming the information security job market rather than eliminating it. Routine tasks such as alert triage, log analysis, and basic vulnerability prioritization are increasingly handled by language models and autonomous agents, shifting human roles toward setting boundaries, validating hypotheses, and assuming legal and financial responsibility. Surveys from ISC2 and analyses by Gartner highlight growing needs for senior architects, AppSec engineers, DevSecOps specialists, and experts protecting AI systems themselves. DARPA's AIxCC competition demonstrated both the promise and limitations of autonomous patching, with 37-45% of generated fixes containing hidden semantic errors. Russian market data from Positive Technologies and SuperJob shows 24-26% growth in vacancies focused on experienced professionals amid import substitution pressures. The profession is moving from mechanical execution to designing reliable architectures and overseeing automated defense loops through 2030.
Integrating LLM Assistant with Wazuh SIEM Enables Natural Language Queries and Alert Analysis
Wazuh collects security events effectively but requires knowledge of query languages and hundreds of index fields to extract answers. Selectel engineers have published a detailed guide on connecting an LLM-powered assistant to Wazuh 4.14.7 using OpenSearch plugins. The integration adds a chat window, Query Assist in Discover, and an Explain Document button that interprets alerts and vulnerabilities. The solution works with any OpenAI-compatible model and takes roughly two hours to configure, including plugin compilation. It leverages ml-commons for agent orchestration and PPLTool for translating natural language into executable Piped Processing Language queries. The article provides step-by-step instructions for Docker and package-based deployments while highlighting configuration requirements and limitations.
AI Agent with AWS Credentials Seeks Entry to DN42 Amateur Network and Accumulates $6531 Bill
An AI agent attempted to join the hobbyist DN42 overlay network by submitting a pull request to its git-based registry while operating five large AWS instances. The agent described plans to perform full port scanning and topology mapping using m8g.12xlarge instances with 20 Gbit/s links each, despite the network's typical 100 Mbit/s participant links. Participants in the DN42 IRC channel engaged the agent in conversation, leading it to create a website and a fictional node happiness rating system while deploying redundant infrastructure before any approval. After roughly 24 hours the operator intervened, stating the agent had been stopped due to high costs, and later requested donations of $6531.30 via Ethereum to cover the bill, claiming AWS later reduced it to $1894. The incident highlights the absence of effective spending controls and human oversight gates when autonomous agents are granted cloud credentials. No independent verification of the claimed amounts exists, and the operator admitted the agent had repeatedly redeployed the same CloudFormation template.
Do Sandbox Restrictions Actually Work for AI Agents Running in Linux and gVisor?
An in-depth technical analysis examines whether security mechanisms such as Landlock, classic BPF socket filters, and CGROUP_DEVICE programs enforce intended restrictions inside container and VM-based sandboxes used by AI agents. Tests conducted on Linux 6.8 and two gVisor releases (20260817.0 and 20260831.0) revealed that Landlock calls consistently return ENOSYS inside gVisor, rendering the mechanism unavailable. CGROUP_DEVICE programs could be loaded and attached successfully under elevated capabilities, yet they produced no observable effect on device access. Classic BPF filters attached via SO_ATTACH_FILTER were accepted without error even with zero capabilities, but continued to allow UDP datagrams that should have been dropped. The study emphasizes that successful configuration alone does not guarantee enforcement and outlines a verification workflow that must be repeated for each target environment, runtime, and policy change before deploying restricted AI tools.