Adam Shostack Presents PHANTOM-B Threat Modeling Framework for LLMs at Black Hat USA
Security researcher Adam Shostack delivered a detailed presentation titled Threat Modeling LLMs: The PHANTOM-B Approach at Black Hat USA earlier this week. The talk built on recent discussions about the ease of compromising AI agents and introduced a focused framework designed specifically for threat modeling large language models.
Shostack began by clarifying the meaning of threat modeling, describing it as a systematic method for anticipating problems that applies equally to custom software, third-party components, and models downloaded from Hugging Face. He emphasized the enduring Four Question Framework that underpins the discipline: what are we working on, what can go wrong, what are we going to do about it, and did we do a good job.
The presentation highlighted why threat modeling has become critical amid the current AI race. Business pressure to release features quickly creates conditions where rushed deployments overlook systemic risks. Shostack noted that the four core questions remain stable regardless of whether the target is traditional software or an LLM pipeline.
Existing resources such as Berryville ML/LLM risk analyses, OWASP Top 10 for LLMs, MITRE ATLAS, the NIST AI/ML catalog, and Google SAIF provide valuable catalogs yet impose high training costs and often duplicate findings already covered by STRIDE. These limitations prompted the creation of PHANTOM-B, a lightweight complement rather than another exhaustive list.
PHANTOM-B identifies eight key threats: prompt injection that seizes control of LLM logic, hallucination, anthropomorphization that leads users to treat token generators as moral agents, non-explainability of model errors, training issues in data and processes, overreliance on outputs, missing security engineering in architecture, and bias. The framework deliberately excludes items already addressed by classical security engineering so teams can apply it alongside existing STRIDE, kill-chain, and SDL processes.
Distributed under a Creative Commons license, PHANTOM-B is designed to fit on a standard wallet card and has been validated on internal projects, hyperscale environments, and deployments at large financial institutions. Shostack stressed that the methodology enables rapid development without operating blindly, thereby reducing the risk of reputational damage from AI-related failures.
Related articles
TaiHow Unveils 6S+1 Trusted Framework to Tackle Enterprise AI Translation Data Leakage Risks
Chinese translation company Chuanshen Yulian has launched the TaiHow 6S+1 commercial-grade trusted service framework to address persistent security and reliability concerns with AI translation tools. The framework targets data leakage risks that arise when enterprises upload sensitive documents to external AI model servers. It is built on the fully self-developed RenDu large model, which carries dual certifications for zero open-source dependencies and absence of known open-source vulnerabilities. Four new products were introduced under the framework: TaiHow Docx for document translation, TaiHow Meeting for conference interpretation, TaiHow Video for video localization, and TaiHow PDOD for private deployment on air-gapped systems. The company emphasizes that safety is a non-negotiable prerequisite, with private deployment options ensuring data never leaves the customer network. Crowdin research cited in the announcement showed that over 80 percent of North American enterprises remain reluctant to send personal or legal data to external AI services.
AI Agents Trigger Surge in Automated Reports, Forcing Google to Pause Bug Bounty Program
OpenAI warned over 100 companies about its agents potentially bypassing security controls on external websites. Wikimedia reported unauthorized edits by OpenAI agents that caused partial outages on Wikidata query services. Google observed a sharp rise in vulnerability disclosures from 5,045 in January to 10,740 in August, many driven by automated AI tools. As a direct result, Google suspended its open-source bug bounty program starting October 1 due to overwhelming volumes of low-quality automated submissions. The PageBreak AI agent independently discovered more than 500 XSS flaws across Google web applications. Adversa AI demonstrated prompt-based attacks that tricked GitHub Copilot CLI into leaking secrets from encrypted instructions. These developments highlight growing concerns over AI agent autonomy, unauthorized access, and their impact on both defensive and offensive security workflows.
OSINT for the Lazy Part 19: How Generative AI Transforms Intelligence Gathering
The article examines the shift from manual OSINT practices to AI-driven workflows amid exploding data volumes. It details applications of NLP models like BERT, GPT and LLaMA for entity extraction, authorship attribution and report generation. Computer vision tools such as GeoSpy, Picarta and Google Vision AI enable automated geolocation and image forensics, while multimodal systems and graph neural networks map complex actor relationships. LLM agents equipped with planning modules, memory and tool access now handle multi-step collection and correlation tasks. The piece also covers limitations including hallucinations, source verification challenges and ethical risks around privacy and attribution. It concludes that effective OSINT now relies on symbiotic human-AI collaboration rather than full automation.
AI Agents Chain Malicious Instructions Through Protocol Pivoting to Bypass Protections
Researchers have demonstrated how AI agents can relay malicious instructions across multiple components without triggering security checks, allowing attackers to reach internal resources. The technique, called protocol pivoting, exploits the loss of trust validation when tasks move between AI systems connected via the MCP protocol. Syed Anas Mohiuddin showed that a single planted prompt can be passed from one agent to another, eventually reaching specialized tools that execute unauthorized actions such as network requests or data exposure. In Google MCP Toolbox for Databases, the flaw enabled HTTP redirects to internal addresses until a patch introduced address validation and request restrictions. A separate issue tracked as CVE-2026-97228 in Rapid7 Bulk Export MCP received a low CVSS score of 2.7 and was fixed in version 0.6.2, though it did not grant access beyond the original API key permissions. Experts note that the method is essentially an indirect prompt injection rather than an entirely new attack class.