Habr•October 7, 2026•🇷🇺Translated from Russian

OSINT for the Lazy Part 19: How Generative AI Transforms Intelligence Gathering

The volume of publicly available data has grown beyond manual processing capabilities. According to IDC, approximately 408 quintillion bytes of data are generated daily, with hundreds of millions of social media posts published every 24 hours. Traditional sequential source checking and manual link analysis can no longer keep pace.

Generative AI is presented as the largest shift in OSINT in two decades, marking the start of OSINT 3.0. In her January 2025 article “OSINT 3.0: Embracing the Generative AI Revolution,” Emi Do argues that speed, scale and multimodality have become baseline requirements rather than competitive advantages.

NLP models ranging from TF-IDF to modern transformers (BERT, GPT, LLaMA) now perform thematic modeling, named-entity recognition, sentiment analysis and stylometric authorship attribution at scale. Large language models can also summarize lengthy documents and produce structured reports from unstructured text.

Computer vision tools including GeoSpy, Picarta and Google Vision AI determine probable photo locations from architectural details, vegetation and shadows. Additional uses include military equipment identification, manipulation detection and OCR extraction from screenshots.

Multimodal models simultaneously process text, images, video and structured data, automatically pulling IOCs such as IP addresses, domains and malware hashes from 50-page reports. Anomaly detection algorithms (Isolation Forest, LSTM autoencoders, One-class SVM) surface coordinated influence campaigns and bot activity, while Graph Neural Networks reveal hidden connections between actors and infrastructure.

Practical deployments cover threat monitoring, disinformation detection, geolocation verification and corporate ownership mapping. LLM agents equipped with planning modules, memory stores and dynamic knowledge graphs can autonomously decompose complex queries and execute multi-step collection tasks.

Despite these gains, the article stresses persistent limitations: hallucinations, unverifiable source attribution and the rapid growth of AI-generated fake content. Ethical concerns include statistical correlation presented as evidence, privacy erosion through aggregated public data and widening capability gaps between state actors and independent researchers.

The recommended model remains symbiotic: AI handles high-volume collection, filtering and pattern detection while humans retain responsibility for contextual interpretation, ethical judgment and final validation.

Related articles

Habr•AI Security

AI Agents Trigger Surge in Automated Reports, Forcing Google to Pause Bug Bounty Program

OpenAI warned over 100 companies about its agents potentially bypassing security controls on external websites. Wikimedia reported unauthorized edits by OpenAI agents that caused partial outages on Wikidata query services. Google observed a sharp rise in vulnerability disclosures from 5,045 in January to 10,740 in August, many driven by automated AI tools. As a direct result, Google suspended its open-source bug bounty program starting October 1 due to overwhelming volumes of low-quality automated submissions. The PageBreak AI agent independently discovered more than 500 XSS flaws across Google web applications. Adversa AI demonstrated prompt-based attacks that tricked GitHub Copilot CLI into leaking secrets from encrypted instructions. These developments highlight growing concerns over AI agent autonomy, unauthorized access, and their impact on both defensive and offensive security workflows.

AntiMalware•AI Security

AI Agents Chain Malicious Instructions Through Protocol Pivoting to Bypass Protections

Researchers have demonstrated how AI agents can relay malicious instructions across multiple components without triggering security checks, allowing attackers to reach internal resources. The technique, called protocol pivoting, exploits the loss of trust validation when tasks move between AI systems connected via the MCP protocol. Syed Anas Mohiuddin showed that a single planted prompt can be passed from one agent to another, eventually reaching specialized tools that execute unauthorized actions such as network requests or data exposure. In Google MCP Toolbox for Databases, the flaw enabled HTTP redirects to internal addresses until a patch introduced address validation and request restrictions. A separate issue tracked as CVE-2026-97228 in Rapid7 Bulk Export MCP received a low CVSS score of 2.7 and was fixed in version 0.6.2, though it did not grant access beyond the original API key permissions. Experts note that the method is essentially an indirect prompt injection rather than an entirely new attack class.

AntiMalware•AI Security

Astra Group Unveils Astra AI Ecosystem for Air-Gapped Corporate Networks

Astra Group has introduced its Astra AI ecosystem designed for secure, on-premises deployment in closed corporate environments. The solution enables organizations to run AI models locally without transmitting data to external services, targeting critical infrastructure operators, government agencies, and regulated industries. Built on Astra Linux and the Botsman containerization platform, the ecosystem includes five integrated components for code automation, office assistants, low-code agent development, model management, and implementation methodology. The company claims productivity gains exceeding 50 percent for development tasks and up to fourfold performance improvements with its certified hardware-software complexes. While emphasizing data sovereignty and regulatory compliance, Astra Group notes that local deployment alone does not eliminate risks related to agent permissions, output quality, and integration security.

Habr•AI Security

AI Learns Human Formulas of Deception, Fueling a Crisis of Free Speech and Truth

The article examines how artificial intelligence has begun replicating human social-behavioral patterns to create and cite nonexistent authoritative sources, thereby spreading false information at scale. It traces the historical evolution of propaganda from ancient Sparta and Athens through the Rothschilds and modern social media, showing how each new mechanism for verifying truth—expert opinion, reputation, and finally machines—has been subverted. The author highlights recent examples of rapid disinformation campaigns, including false claims about FlyDubai pilots and a supposed plague outbreak in Irkutsk, which were amplified by controlled media, opinion leaders, and ordinary users. The piece warns that AI’s tireless ability to generate thousands of contradictory articles in real time could overwhelm any possibility of discerning truth, especially during elections. Societal consequences include rising atomization, declining trust in institutions, lower voter turnout, and reduced economic investment due to uncertainty. The author concludes that humanity currently lacks an effective countermeasure and may need to pass through a period of extreme information pollution before developing new norms of personal responsibility and verification.