Habr•September 9, 2026•🇷🇺Translated from Russian

PII-Guard: Open-Source Detector for Personal Data in Russian Text

PII-Guard, an open-source detector for personal data in Russian text, has been released by NLP researcher Andrey Ivanov from the R&D laboratory of red_mad_robot. The system addresses the growing need to protect sensitive information when users and companies send working texts to language models.

Personal data such as names, phone numbers, document numbers and addresses frequently accompany prompts. Companies therefore build masking pipelines that anonymize text before model inference. However, simple masking destroys semantic links, forcing models to answer questions about “**** called ****” without knowing how many people are involved or who called whom.

The new requirement is reversible pseudonymization: real values must be replaced by structured placeholders that models can process naturally, after which the original data is restored in the final answer.

From regex to hybrid detection

The first prototype relied on regular expressions. While telephone numbers, INN, SNILS and card numbers have predictable length and structure, real-world text contains spaces, dashes, typos and extra characters that break rigid patterns. Moreover, a 16-digit sequence could be either a bank card or an order ID, and a 10-digit string could match both INN and passport formats.

Control-sum verification was added next. Bank cards use the Luhn algorithm, INN and SNILS employ weighted sums, and OMS policies again apply Luhn. Additional structural checks cover BIK prefixes and postal-index ranges. These filters reduced false positives, yet arithmetic validation alone cannot distinguish a genuine document from a coincidental numeric string that happens to satisfy the checksum.

Context windows were introduced to look for nearby cue words such as “паспорт”, “ИНН” or “полис”. Positive keywords must lie within type-specific distance limits; negative keywords such as “биржевой” can veto an index match regardless of distance. When multiple candidates compete, the system selects the highest-confidence match after ordering document types from most distinctive (birth certificates, military IDs) to most ambiguous (simple six- or ten-digit sequences).

Adding semantic understanding

Names, addresses and telephone numbers lack rigid formats, so rules cannot locate them reliably. The team therefore annotated their own datasets and fine-tuned a ruBert-base NER model. The model handles semantic entities such as person names and addresses while also capturing documents missed by rules because of typos or irregular spacing.

The final architecture runs two independent streams. The rule-based stream normalizes text, extracts candidates, validates checksums and applies keyword context. The model stream processes the entire text and returns its own spans. An arbitration module merges results: when spans overlap but labels differ, rules win; when only the model finds an entity, its prediction is accepted; non-overlapping spans from both streams are retained.

Pseudonymization and grammatical restoration

Detected entities are replaced by XML-style tags containing type, numeric ID and, for persons, grammatical gender inferred by a Russian morphology library. The tag format was chosen to avoid accidental collisions with ordinary text or other system markup. After the model returns its answer, real values are substituted back, with morphological analysis ensuring correct case for names and addresses.

Evaluation was performed on four public datasets, one of which is the team’s own. On the intersection of entity types supported by all compared systems, PII-Guard achieved the highest micro-F1 scores under both strict span matching and type-overlap criteria. The project, datasets and code are publicly available.

Related articles

Habr•Privacy & Surveillance

CookieTin Extension Manages Partitioned Cookies Across Firefox, Chrome and Edge

Developer Perruer2 has released CookieTin, an open-source browser extension that fully supports partitioned cookies under Firefox Total Cookie Protection and Chrome CHIPS. The tool addresses limitations in older managers like Cookie Quick Manager by correctly retrieving and deleting cookies stored with partitionKey values. It works across Firefox, Chrome and Edge using a single Manifest V3 codebase written in TypeScript and Preact. Key features include accurate cookies.txt export compatible with curl and yt-dlp, protected cookies that survive explicit deletion, and pre-save validation of browser rules for __Host- prefixes and SameSite attributes. E2E tests using Puppeteer verify handling of HttpOnly, partitioned and container cookies in all three browsers.

AntiMalware•Privacy & Surveillance

Kaspersky Premium for macOS Gains App Uninstall Feature to Remove Residual Files

Kaspersky Premium now includes an App Uninstall tool for macOS that locates and deletes leftover files such as caches, cookies, settings, and logs after applications are removed. The feature also identifies duplicate copies of programs and lets users remove all instances or select specific ones while preserving shared components used by other software. Survey data from Kaspersky shows that only 44 percent of macOS users delete unused applications, even though 56 percent regularly clear browser data and 54 percent remove unwanted media files. Residual files can contain sensitive information including account tokens, passwords, IP addresses, event logs, and personal documents, creating privacy risks especially when a device is sold or accessed by unauthorized parties. Deleted files can be restored from the trash or directly within Kaspersky Premium before the application session ends. The company also warns that malicious programs are frequently disguised as legitimate macOS cleaning utilities.

Securitylab•Privacy & Surveillance

Bypassing VPN Detection on iPhone: Detailed Methods to Avoid App Blocks

Many iPhone users encounter apps that detect and block active VPN connections even after switching servers or protocols. The detection often occurs locally on the device by inspecting network interfaces rather than relying solely on external IP addresses. This guide explains how apps identify VPN tunnels through iOS network data and provides practical workarounds including moving the VPN to a router, configuring per-app exclusions, and using web versions of services. It also covers why protocol obfuscation and port changes fail to hide local VPN activity from applications. Additional troubleshooting addresses automatic VPN profiles, ad blockers, and iCloud Private Relay interference. The article emphasizes that no universal toggle exists in iOS to hide an active VPN from all apps.

Habr•Privacy & Surveillance

New Obfuscation Method Dissolves Personal Data Records in Layer of Plausible Variants

A Russian information security researcher has proposed a data protection technique that renders stolen personal records unusable even after full compromise. The approach mixes real data such as phone numbers, emails, passports, addresses, INN and SNILS with vast numbers of semantically valid alternatives. Attackers receive nearly complete information including a 361-character message containing PIN codes and word order, yet lack the secret vector space and reconstruction algorithm required to identify the correct record. Without these components, brute-force attempts produce millions of plausible results with no architectural method to verify accuracy. The method is presented as an alternative to traditional encryption when data must remain accessible yet protected against extraction. A public sandbox is available for testing the approach.