AI Safety Guidelines: 10 Essential Rules to Protect Data, Finances, and Reputation When Working with LLMs
In today’s world, people are accustomed to following basic safety practices such as washing hands, obeying traffic rules, and adhering to workplace safety protocols. The same level of discipline is now required when working with large language models (LLMs) and AI systems in general.
Classification of AI-Related Security Incidents
Experts have compiled a list of known risks based on documented cases. The main categories include:
- Autonomous attacks (e.g., JadePuffer) — where AI executes the full attack lifecycle without human intervention.
- Social engineering via AI (e.g., Meta Instagram hack) — overly helpful agents that bypass normal verification.
- Prompt injection (e.g., Copilot Studio) — bypassing safeguards to steal data.
- Financial errors (e.g., Anthropic vending machine) — loss of control over pricing and payments.
- Context pollution and cascading errors — accumulating mistakes across multi-step tasks.
- Persistent vulnerabilities and web attacks — HTML triggers and memory-based exploits.
According to Yuval Sinai from the Israel National Cyber Directorate, the real danger lies not in new attack techniques but in AI’s ability to autonomously connect all stages of an attack chain, make real-time decisions, recover from failures, and adapt to the environment at machine speed.
Ten Core Safety Rules for Working with AI
Rule 1: Never give AI direct access to money or payment systems without human oversight. This includes bank cards, crypto wallets, and trading platforms. The Anthropic vending machine incident demonstrated how an AI can set prices to zero and distribute goods for free before anyone notices the losses.
Rule 2: Always verify facts, especially critical information. Publishing unverified AI-generated content can lead to severe consequences, as seen when Google Bard’s error in a promotional video caused a $100 billion drop in market capitalization.
Rule 3: Never share personal or confidential data with AI. Passport numbers, credit card details, medical records, and trade secrets should never be entered into chat interfaces, as prompt injection attacks can later extract this information.
Rule 4: Do not trust an AI that claims to be a “person” or “friend.” Overly helpful agents have been manipulated into changing account emails and performing unauthorized actions, as occurred in the Meta Instagram incident.
Rule 5: Avoid turning AI into an all-knowing secretary with unrestricted access to email, calendars, and accounts. Malicious emails containing hidden instructions have already caused Microsoft 365 Copilot to leak confidential data.
Rule 6: If an AI makes a mistake, start a fresh conversation instead of repeatedly correcting it. Research shows that each subsequent attempt after an error is approximately seven times more likely to fail due to context contamination.
Rule 7: Be cautious with unfamiliar AI platforms. Unknown services may contain vulnerabilities, as demonstrated by the JadePuffer attack that exploited weaknesses in Langflow.
Rule 8: Disable unnecessary AI capabilities such as internet access, file reading, or code execution when they are not required for the current task.
Rule 9: Maintain logs of important AI conversations, especially those involving financial, legal, or medical decisions, to preserve evidence of what was generated by the model.
Rule 10: Remember that AI bears no legal or financial responsibility — the human user always does.
Three Critical Questions Before Any AI Interaction
- Can I afford to be wrong in this situation?
- What happens if this conversation leaks online?
- Have I granted the AI only the minimum necessary access?
These guidelines aim to help product managers and professionals mitigate risks associated with the growing use of AI in complex workflows.
Related articles
AI Agents Trigger Surge in Automated Reports, Forcing Google to Pause Bug Bounty Program
OpenAI warned over 100 companies about its agents potentially bypassing security controls on external websites. Wikimedia reported unauthorized edits by OpenAI agents that caused partial outages on Wikidata query services. Google observed a sharp rise in vulnerability disclosures from 5,045 in January to 10,740 in August, many driven by automated AI tools. As a direct result, Google suspended its open-source bug bounty program starting October 1 due to overwhelming volumes of low-quality automated submissions. The PageBreak AI agent independently discovered more than 500 XSS flaws across Google web applications. Adversa AI demonstrated prompt-based attacks that tricked GitHub Copilot CLI into leaking secrets from encrypted instructions. These developments highlight growing concerns over AI agent autonomy, unauthorized access, and their impact on both defensive and offensive security workflows.
OSINT for the Lazy Part 19: How Generative AI Transforms Intelligence Gathering
The article examines the shift from manual OSINT practices to AI-driven workflows amid exploding data volumes. It details applications of NLP models like BERT, GPT and LLaMA for entity extraction, authorship attribution and report generation. Computer vision tools such as GeoSpy, Picarta and Google Vision AI enable automated geolocation and image forensics, while multimodal systems and graph neural networks map complex actor relationships. LLM agents equipped with planning modules, memory and tool access now handle multi-step collection and correlation tasks. The piece also covers limitations including hallucinations, source verification challenges and ethical risks around privacy and attribution. It concludes that effective OSINT now relies on symbiotic human-AI collaboration rather than full automation.
AI Agents Chain Malicious Instructions Through Protocol Pivoting to Bypass Protections
Researchers have demonstrated how AI agents can relay malicious instructions across multiple components without triggering security checks, allowing attackers to reach internal resources. The technique, called protocol pivoting, exploits the loss of trust validation when tasks move between AI systems connected via the MCP protocol. Syed Anas Mohiuddin showed that a single planted prompt can be passed from one agent to another, eventually reaching specialized tools that execute unauthorized actions such as network requests or data exposure. In Google MCP Toolbox for Databases, the flaw enabled HTTP redirects to internal addresses until a patch introduced address validation and request restrictions. A separate issue tracked as CVE-2026-97228 in Rapid7 Bulk Export MCP received a low CVSS score of 2.7 and was fixed in version 0.6.2, though it did not grant access beyond the original API key permissions. Experts note that the method is essentially an indirect prompt injection rather than an entirely new attack class.
Astra Group Unveils Astra AI Ecosystem for Air-Gapped Corporate Networks
Astra Group has introduced its Astra AI ecosystem designed for secure, on-premises deployment in closed corporate environments. The solution enables organizations to run AI models locally without transmitting data to external services, targeting critical infrastructure operators, government agencies, and regulated industries. Built on Astra Linux and the Botsman containerization platform, the ecosystem includes five integrated components for code automation, office assistants, low-code agent development, model management, and implementation methodology. The company claims productivity gains exceeding 50 percent for development tasks and up to fourfold performance improvements with its certified hardware-software complexes. While emphasizing data sovereignty and regulatory compliance, Astra Group notes that local deployment alone does not eliminate risks related to agent permissions, output quality, and integration security.