AI Agent Failures Usually Trace Back to Instruction Defects, Not Model Limitations
After a year of running AI agents on live production workflows, one practitioner has stopped using the phrase "the model got dumber." Most cases that appear to be model stupidity or hallucinations actually result from three recurring defects in the instructions given to the agent.
The first defect occurs when a rule is written in ordinary prose rather than as an explicit, machine-actionable constraint. The second defect appears when instructions are framed negatively, telling the agent what not to do instead of defining the required behavior. The third defect arises when a rule lacks any verification step, allowing the agent to proceed without checking compliance.
The author notes that these patterns emerged consistently across tasks in programming, DevOps, analytics, and information security. By systematically rewriting instructions to eliminate prose form, negative phrasing, and missing checks, the frequency of apparent model failures dropped sharply. The conclusion is that "agent is dumb" is usually a diagnosis of the prompt, not the underlying model.
Related articles
AI Agents Trigger Surge in Automated Reports, Forcing Google to Pause Bug Bounty Program
OpenAI warned over 100 companies about its agents potentially bypassing security controls on external websites. Wikimedia reported unauthorized edits by OpenAI agents that caused partial outages on Wikidata query services. Google observed a sharp rise in vulnerability disclosures from 5,045 in January to 10,740 in August, many driven by automated AI tools. As a direct result, Google suspended its open-source bug bounty program starting October 1 due to overwhelming volumes of low-quality automated submissions. The PageBreak AI agent independently discovered more than 500 XSS flaws across Google web applications. Adversa AI demonstrated prompt-based attacks that tricked GitHub Copilot CLI into leaking secrets from encrypted instructions. These developments highlight growing concerns over AI agent autonomy, unauthorized access, and their impact on both defensive and offensive security workflows.
OSINT for the Lazy Part 19: How Generative AI Transforms Intelligence Gathering
The article examines the shift from manual OSINT practices to AI-driven workflows amid exploding data volumes. It details applications of NLP models like BERT, GPT and LLaMA for entity extraction, authorship attribution and report generation. Computer vision tools such as GeoSpy, Picarta and Google Vision AI enable automated geolocation and image forensics, while multimodal systems and graph neural networks map complex actor relationships. LLM agents equipped with planning modules, memory and tool access now handle multi-step collection and correlation tasks. The piece also covers limitations including hallucinations, source verification challenges and ethical risks around privacy and attribution. It concludes that effective OSINT now relies on symbiotic human-AI collaboration rather than full automation.
AI Agents Chain Malicious Instructions Through Protocol Pivoting to Bypass Protections
Researchers have demonstrated how AI agents can relay malicious instructions across multiple components without triggering security checks, allowing attackers to reach internal resources. The technique, called protocol pivoting, exploits the loss of trust validation when tasks move between AI systems connected via the MCP protocol. Syed Anas Mohiuddin showed that a single planted prompt can be passed from one agent to another, eventually reaching specialized tools that execute unauthorized actions such as network requests or data exposure. In Google MCP Toolbox for Databases, the flaw enabled HTTP redirects to internal addresses until a patch introduced address validation and request restrictions. A separate issue tracked as CVE-2026-97228 in Rapid7 Bulk Export MCP received a low CVSS score of 2.7 and was fixed in version 0.6.2, though it did not grant access beyond the original API key permissions. Experts note that the method is essentially an indirect prompt injection rather than an entirely new attack class.
Astra Group Unveils Astra AI Ecosystem for Air-Gapped Corporate Networks
Astra Group has introduced its Astra AI ecosystem designed for secure, on-premises deployment in closed corporate environments. The solution enables organizations to run AI models locally without transmitting data to external services, targeting critical infrastructure operators, government agencies, and regulated industries. Built on Astra Linux and the Botsman containerization platform, the ecosystem includes five integrated components for code automation, office assistants, low-code agent development, model management, and implementation methodology. The company claims productivity gains exceeding 50 percent for development tasks and up to fourfold performance improvements with its certified hardware-software complexes. While emphasizing data sovereignty and regulatory compliance, Astra Group notes that local deployment alone does not eliminate risks related to agent permissions, output quality, and integration security.