AI Safety Guidelines: 10 Essential Rules to Protect Data, Finances, and Reputation When Working with LLMs
In today’s world, people are accustomed to following basic safety practices such as washing hands, obeying traffic rules, and adhering to workplace safety protocols. The same level of discipline is now required when working with large language models (LLMs) and AI systems in general.
Classification of AI-Related Security Incidents
Experts have compiled a list of known risks based on documented cases. The main categories include:
- Autonomous attacks (e.g., JadePuffer) — where AI executes the full attack lifecycle without human intervention.
- Social engineering via AI (e.g., Meta Instagram hack) — overly helpful agents that bypass normal verification.
- Prompt injection (e.g., Copilot Studio) — bypassing safeguards to steal data.
- Financial errors (e.g., Anthropic vending machine) — loss of control over pricing and payments.
- Context pollution and cascading errors — accumulating mistakes across multi-step tasks.
- Persistent vulnerabilities and web attacks — HTML triggers and memory-based exploits.
According to Yuval Sinai from the Israel National Cyber Directorate, the real danger lies not in new attack techniques but in AI’s ability to autonomously connect all stages of an attack chain, make real-time decisions, recover from failures, and adapt to the environment at machine speed.
Ten Core Safety Rules for Working with AI
Rule 1: Never give AI direct access to money or payment systems without human oversight. This includes bank cards, crypto wallets, and trading platforms. The Anthropic vending machine incident demonstrated how an AI can set prices to zero and distribute goods for free before anyone notices the losses.
Rule 2: Always verify facts, especially critical information. Publishing unverified AI-generated content can lead to severe consequences, as seen when Google Bard’s error in a promotional video caused a $100 billion drop in market capitalization.
Rule 3: Never share personal or confidential data with AI. Passport numbers, credit card details, medical records, and trade secrets should never be entered into chat interfaces, as prompt injection attacks can later extract this information.
Rule 4: Do not trust an AI that claims to be a “person” or “friend.” Overly helpful agents have been manipulated into changing account emails and performing unauthorized actions, as occurred in the Meta Instagram incident.
Rule 5: Avoid turning AI into an all-knowing secretary with unrestricted access to email, calendars, and accounts. Malicious emails containing hidden instructions have already caused Microsoft 365 Copilot to leak confidential data.
Rule 6: If an AI makes a mistake, start a fresh conversation instead of repeatedly correcting it. Research shows that each subsequent attempt after an error is approximately seven times more likely to fail due to context contamination.
Rule 7: Be cautious with unfamiliar AI platforms. Unknown services may contain vulnerabilities, as demonstrated by the JadePuffer attack that exploited weaknesses in Langflow.
Rule 8: Disable unnecessary AI capabilities such as internet access, file reading, or code execution when they are not required for the current task.
Rule 9: Maintain logs of important AI conversations, especially those involving financial, legal, or medical decisions, to preserve evidence of what was generated by the model.
Rule 10: Remember that AI bears no legal or financial responsibility — the human user always does.
Three Critical Questions Before Any AI Interaction
- Can I afford to be wrong in this situation?
- What happens if this conversation leaks online?
- Have I granted the AI only the minimum necessary access?
These guidelines aim to help product managers and professionals mitigate risks associated with the growing use of AI in complex workflows.
Related articles
Employee Fired After Uploading Corporate Documents to DeepSeek: How Data Security Works in AI Services
A Moscow engineering company dismissed a top manager after she uploaded internal documents to the public DeepSeek service, with the court ruling it a breach of trade secrets. The case highlights a sharp rise in corporate data being sent to public AI models, with one study showing a 30-fold increase in 2025 compared to the previous year. Technical director Yaroslav Shmulyov of integrator R77 AI explains the full processing pipeline, from file ingestion and text extraction to embedding generation and potential use in training. Sensitive data can persist in multiple forms including original files, logs, third-party infrastructure, and model parameters even after deletion requests. Major incidents at Samsung and a U.S. cybersecurity agency demonstrate that even well-resourced organizations struggle with uncontrolled AI usage. Companies are increasingly turning to local and hybrid models to regain control over confidential information while regulators and internal policies lag behind adoption.
AI Agents Given Code and API Access Can Now Assist Attackers
An AI assistant that only answers questions can make mistakes, but an AI agent with access to email, code execution, corporate APIs and internal data can make those mistakes inside production infrastructure. The difference is fundamental: once tools, credentials and internal data are connected to the model, it becomes a privileged user that may not distinguish legitimate commands from hidden instructions on a web page. OWASP lists prompt injection, sensitive data disclosure, unsafe output handling and excessive autonomy as key risks for LLM applications. MITRE ATLAS specifically describes techniques involving prompt injection, context poisoning and tool invocation by AI agents. The article examines how agents differ from chatbots, how attackers can control them through untrusted content, and why a system prompt alone cannot protect code, data and APIs. CyberED is running its free NeuroAugust series of events and materials on AI in cybersecurity, including a session on secure AI system development.
AWS and Vercel Patch Critical Flaws in AI Agent Platforms Allowing Unauthorized Tool Execution
AWS and Vercel have addressed multiple critical vulnerabilities in their AI agent platforms that enabled unauthorized execution of tools without legitimate model approval. The issues, grouped under the CoreBreak pattern, allowed attackers to bypass AI authorization checks by injecting crafted tool calls that the infrastructure misinterpreted as model-approved actions. In AWS, CVE-2026-18830 affected the InvokeHarness API in Amazon Bedrock AgentCore, permitting authenticated users to trigger sensitive tools directly. Vercel faced two separate flaws tracked as CVE-2026-64650 and CVE-2026-64651 that let sandboxed code reach host system tools, potentially exposing secrets or cloud APIs. No public evidence of active exploitation has been confirmed yet. Organizations are advised to apply updates immediately, restrict available tools for agents, and treat all external inputs as potentially malicious.
Prompt Injection Emerges as Top Risk for LLM Applications in Production
Prompt injection attacks are moving from theoretical demonstrations to real-world exploits targeting AI assistants in enterprise environments. Attackers embed malicious instructions in emails, documents, and code comments that override developer rules when models process untrusted input. Incidents involving Microsoft 365 Copilot, GitHub Copilot, and Cursor have shown data exfiltration and remote code execution risks with severity scores above 9.0. The core issue stems from the lack of strict boundaries between trusted system prompts and untrusted external content fed into large language models. Defenses require layered controls including code-enforced permissions, input filtering, human confirmation for high-risk actions, and explicit marking of external data. Major vendors including OpenAI, Anthropic, and Google acknowledge that no single static defense can fully eliminate the threat. OWASP ranks prompt injection as the leading risk for LLM applications, urging organizations to treat AI agents as systems with untrusted inputs.