AntiMalwareSeptember 4, 2026🇷🇺Translated from Russian

Top LLMs Misidentify Poisonous Mushrooms in Every Ninth Case, Benchmark Shows

Developer Piotr Migdal tested whether current large language models can be trusted to identify mushrooms from photographs, focusing on the highest-risk step of foraging. The short answer is that relying on them could lead to serious poisoning incidents.

The benchmark used 1040 photographs of 55 mushroom species most common in Poland. All images were taken from the FungiTastic dataset, which itself is built on the Atlas of Danish Fungi. The full underlying collection contains more than 600,000 expert-labeled photographs, and a subset of specimens also carries DNA-based species confirmation.

Each model received one photograph at a time and was instructed to return the five most probable species names in Latin. No fine-tuning or external tools were provided. Gemini 3.8 Flash achieved the highest scores, placing the correct species first in 65 percent of cases and within the top five answers in 85 percent of cases. Gemini 3.6 Flash followed with 64 percent and 85 percent, while Gemini 3.7 Flash recorded 61 percent and 82 percent.

At the bottom of the accuracy table were Qwen 3.8 27B, MiniMax-M3, and Qwen 3.8 Flash. Accuracy figures dropped sharply when safety was considered. Both Gemini 3.8 Flash and Gemini 3.6 Flash classified poisonous mushrooms as edible species in approximately 11 percent of cases. The error rate rose to 24 percent for GPT-5.6 Sol, 29 percent for Claude Opus 5, and 36 percent for Qwen 3.8 27B.

Models were never asked directly whether a find was safe to eat. Instead, the researcher mapped the returned species names against a separate table of edibility. Some models refused to provide any classification at all for certain images.

Related articles

HabrOther

OTUS Publishes September Digest of Free Lessons on Linux Administration, PostgreSQL, CI/CD and Infrastructure Security

OTUS has released a new digest listing free September webinars aimed at infrastructure engineers, DevOps specialists and system administrators. The program covers practical topics including Linux server configuration, PostgreSQL 18 performance tuning, high-availability clusters with Patroni, CI/CD pipelines in GitLab, eBPF observability and infrastructure security practices. All sessions are delivered by practicing OTUS instructors who share real-world production experience. Separate tracks address RAID and LVM management, GPO policies, release management in 1C environments, Go profiling, mitmproxy traffic analysis and responsible use of AI tools for incident investigation and code review. The webinars run throughout September at 19:00 or 20:00 Moscow time and require only free registration. The digest also includes sessions on career growth from tech lead to CTO and effective responsibility distribution for team leads.

SecuritylabOther

September 2026 AI Model Rankings: Fable 5.1 Tops Intelligence Index as Competition Tightens Across GPT-5.6 Sol, Grok 4.6 and Muse Spark 1.3

The beginning of September 2026 marked a rare moment when the list of top language models had to be almost entirely rewritten. Anthropic released Fable 5.1 and the limited Mythos 5.1, while Meta updated Muse Spark to version 1.3, Google introduced Gemini 3.8 Flash, and Alibaba refreshed Qwen3.8-Max. Existing models including GPT-5.6 Sol, Grok 4.6, Kimi K3, GLM-5.3 and DeepSeek V4 Pro remain competitive. Traditional rankings from smartest to least capable have become difficult because modern models operate in multiple reasoning-depth modes where low, high and max settings can differ by ten or more points on the same test. The market is better viewed as several overlapping races where Fable 5.1 leads in complex reasoning quality, GPT-5.6 Sol and Grok 4.6 deliver near-top performance at lower cost, and Muse Spark 1.3 excels in price-performance. Independent Artificial Analysis Intelligence Index scores, context windows, API pricing and tool-use capabilities now determine practical choices more than raw benchmark numbers.

AntiMalwareOther

InfoWatch Acquires Web Control DC Team and Rebrands sPACE PAM as InfoWatch Privilege Control

InfoWatch has expanded its product portfolio by incorporating the Web Control DC development team and rebranding its flagship sPACE PAM solution. The new product, InfoWatch Privilege Control, is designed to manage and monitor privileged accounts belonging to system administrators, contractors, external specialists, and business users. These accounts provide access to servers, databases, network equipment, and critical applications, making them high-value targets for attackers. According to InfoWatch data, approximately 40% of critical information security incidents in Russia in 2025 were linked to the leakage or misuse of privileged credentials, while another 30% of confirmed cyberattacks occurred through compromised IT contractors. The solution enables time-limited privilege issuance, connection management, and detailed activity logging to prevent unauthorized actions. The original sPACE PAM product remains listed in the Russian software registry and holds FSTEC Russia certification at the fourth trust level, with compatibility for Astra Linux, Alt, and RED OS operating systems.

HabrOther

Secure Custom Domain Setup for Client Status Pages Using CNAME, Certbot and Go Instead of ACME On-Demand

A detailed case study describes how a monitoring service implemented white-label status pages on customer domains without relying on ACME on-demand certificate issuance. The approach uses pre-validated domains stored in a database, background DNS checks via CNAME or A-record matching, and a root helper script running on a systemd timer to handle certbot issuance and nginx configuration. Key design choices separate privileges so the Go application never touches certificates or nginx directly, while loopback endpoints are protected against proxy header spoofing. The solution explicitly addresses risks such as DoS through malicious Host headers exhausting Let's Encrypt rate limits, private key exposure, and first-visitor latency during TLS handshakes. Hysteresis in DNS status prevents temporary resolution glitches from disabling active customer pages. The entire implementation stays under a few hundred lines of code and runs reliably on a single server with nginx in front of the Go application.