Unicode Tricks Let Malicious Python Code Bypass Code Review
A subtle but powerful class of attacks allows source code to pass human review while executing completely different logic. The root cause lies in the gap between how fonts render Unicode characters and how the Python interpreter reads the underlying bytes.
The first technique uses homoglyphs. An attacker can declare is_admin = False followed by a second assignment that looks identical but actually uses the Ukrainian letter і (U+0456) instead of Latin i (U+0069). The review sees a single variable being set to True, yet the interpreter creates two separate identifiers. The same substitution works for function names such as validate_token, causing the wrong implementation to run in production.
Python normalizes identifiers with NFKC, which collapses some visually similar characters but leaves Cyrillic letters untouched. Consequently, the attack remains effective in real codebases.
The second technique relies on bidirectional override characters (RIGHT-TO-LEFT OVERRIDE U+202E and related controls). These characters are required for Arabic and Hebrew text, yet they also allow an attacker to reorder source lines visually. A comment that appears harmless on screen can actually contain executable statements once the compiler ignores the display order. The technique became widely known in 2021 as the Trojan Source attack and affects nearly every mainstream programming language.
The third, simpler method inserts zero-width characters (U+200B, U+2060, etc.) inside string literals. The string 'password' has length nine and will never equal the expected value, yet the difference is invisible to the naked eye and defeats both manual inspection and simple grep searches.
A compact detection script walks the source with Python’s tokenize module and separately scans for bidirectional and zero-width characters. Checking only identifier tokens prevents false positives from comments and string data. The same logic is available as built-in rules in ruff and flake8, while gitleaks can block commits containing the dangerous characters.
Teams are advised to run the check across existing repositories, integrate it into CI pipelines, and treat any intentional insertion of these characters as a serious red flag during incident review.
Related articles
CrowdSec Confirms Theft of Source Code from Roughly 300 GitHub Repositories via TanStack Supply Chain Attack
French cybersecurity firm CrowdSec has confirmed that attackers stole source code from approximately 300 GitHub repositories, including around 170 private ones. The breach occurred in May 2026 through a compromised TanStack component that exfiltrated an API key with read access to the private codebase. The stolen material included code for the company's SaaS console, AWS procedures, connectors, and automation tools, while the remaining repositories contained already-public open source code. No customer data, passwords, organization details, tokens, or other secrets were included in the leak, and all potentially affected credentials were immediately rotated. CrowdSec stated that the code is tightly integrated with internal systems and has largely changed over the past four months, reducing its usefulness outside the company's environment. The SaaS service code undergoes regular audits, and the company sees no immediate threat from the exposure while the investigation continues.
Dependency Confusion Attacks Let Attackers Hijack Internal Library Names in Corporate Builds
A widespread supply chain risk allows attackers to publish packages with internal company names on public registries such as npm and PyPI, causing build systems to pull malicious versions instead of internal ones. The attack works because package managers treat multiple registries as a single list and select the highest version number, with no inherent priority for internal sources. Researcher Alex Birsan demonstrated the technique in February 2021 by registering names harvested from open repositories and error messages, successfully injecting packages into builds at Microsoft, Apple, PayPal, Shopify, Netflix, Tesla and Uber. The malicious code executes during installation because setup scripts and lifecycle hooks run with the privileges of the build agent, exposing environment variables, tokens and internal network access. Mitigation requires a single internal proxy repository that never mixes public responses for internal package names, scoped namespaces bound to private registries, lock files with content hashes, and disabling install scripts where possible. The technique remains effective against any organization that lists both internal and public registries in its build configuration.
NEOMSA ESB Release Strengthens Supply Chain Security Through SBOM and Dependency Hardening
Neoflex has released a new version of its NEOMSA ESB integration platform with a primary focus on cleaning up the software bill of materials and eliminating critical and high-severity vulnerabilities. The team automated SBOM generation using CycloneDX, ran SCA scans with Grype and OWASP Dependency-Check, and performed SAST and secret scanning across all build pipelines. Instead of blindly updating to the latest versions, engineers applied minimal fixed versions for each advisory while handling complex cases involving transitive dependencies, locked files, and deprecated build tools. The effort reduced the total package count from 5,872 to 1,683 after migrating the frontend build to Vite in Camel Karavan 4.18. Remaining medium and low findings were tracked in DefectDojo with clear remediation timelines. The changes deliver measurable risk reduction for on-premise deployments in critical infrastructure and financial organizations.
CodeScoring Launches CodeScoring.Save Artifact Repository for Secure Enterprise Development
CodeScoring has introduced its own artifact storage solution called CodeScoring.Save, designed to handle packages, libraries, container images, and other software components used in development. The product targets corporate users of any size seeking a predictable and resilient repository that integrates security checks directly into storage and distribution workflows. Built in Go for modern Kubernetes environments, Save supports multiple package formats including Maven, npm, NuGet, PyPI, Go Modules, Docker/OCI, DEB, and RPM while providing proxy access to external repositories. It features role-based access, auditing, independent scaling of compute and storage layers, and native integration with CodeScoring.OSA to surface vulnerability data inside the repository itself. The company positions Save as a standalone local deployment option that can operate independently or alongside its existing OSA Proxy module to block malicious components at the repository level. Future plans include support for AI models as artifacts, starting with storage and distribution for ecosystems such as Hugging Face, along with certification for Russian secure development requirements.