Physicians Fail to Detect Flawed AI Treatment Recommendations Even When Evidence Contradicts Them, New Study Finds
A new study has demonstrated that physicians often fail to notice when artificial intelligence systems provide flawed treatment recommendations, even when clear evidence of the error is presented directly to them. The research highlights a persistent problem known as automation bias, in which humans tend to place excessive trust in algorithmic outputs that appear objective and computationally derived.
In the experiment, 223 physicians anonymously participated in a series of online scenarios. They were asked to imagine treating patients with a rare disease and to decide whether to administer an experimental therapy whose effectiveness had not yet been confirmed. Before making their decisions, participants received recommendations from an AI system that divided patients into groups based on predicted likelihood of benefit. After selecting patients for treatment, the doctors reviewed actual treatment outcomes and were asked to evaluate the reliability of the AI suggestions.
Researchers deliberately designed the experiments so that the AI predictions would conflict with real results. In the first series, the therapy produced a uniform moderate effect across all patients, yet the algorithm created the false impression that certain groups benefited more than others. In the second series, the treatment was completely ineffective for everyone, but the majority of participants still did not conclude that the therapy offered no benefit at all.
Even after reviewing the outcome data, most doctors continued to view the AI as a reliable source of guidance. They found it difficult to change their initial opinions when new evidence contradicted the algorithm’s conclusions. This effect persisted regardless of the physicians’ level of medical experience, showing that human oversight alone does not automatically catch AI mistakes.
The study carries important implications for the growing use of AI tools in medicine. Such systems are already being deployed or tested to assess complication risks, select treatment strategies, and identify patients requiring closer monitoring. While these tools are intended only as decision-support aids, the research shows that flawed recommendations can still influence clinical choices, potentially leading doctors to withhold effective treatments or administer ineffective ones.
Because the study was conducted in a simulated environment rather than real clinical settings, researchers could precisely control conditions and know exactly where the algorithm was wrong. The authors conclude that simply warning users about possible errors is insufficient. They recommend implementing structured procedures such as requiring an independent evaluation of each case before the AI suggestion is shown, mandating written explanations for agreement with the algorithm, and conducting regular reviews of cases where AI predictions were not confirmed by actual outcomes.
Related articles
Inside the Fortress: Why Perimeter Security Tools Fall Short and How Microsegmentation Protects Networks Internally
Companies invest heavily in perimeter defenses such as firewalls and intrusion detection systems, yet these measures no longer guarantee safety as attackers increasingly operate from within networks. Traditional L2 domains leave virtual machines unisolated, enabling traffic interception, lateral movement, and malware spread similar to an apartment building with poor soundproofing. Microsegmentation powered by SDN divides VLANs into isolated microsegments down to individual VM ports, enforcing granular policies based on ports, IP addresses, and protocols. This approach implements Zero Trust by placing virtual packet filters directly at VM network interfaces on the hypervisor, independent of guest OS actions. Performance remains high because filtering runs on powerful virtualization servers, and scaling occurs naturally as additional hypervisors absorb new workloads without extra configuration. A real-world case from the oil and gas sector shows one customer creating up to 5,000 new microsegmentation rules per week via open REST API. The technology complements rather than replaces perimeter firewalls, delivering both strict internal controls and operational agility.
Good Bear 1.0 Released: Firefox-Based Browser with Isolated Russian PKI Trust Container
Good Bear 1.0 is a Russian-language browser built on Firefox 156.0 that provides an isolated container for handling Russian PKI certificates without mixing trust contexts or user data with the standard browsing session. The release includes .deb packages for Ubuntu 24.04 LTS amd64 and Windows x64 installers, using Mozilla Public License 2.0 and reproducible build processes from pinned Firefox sources. Instead of globally importing root certificates, the browser performs secondary chain validation only inside a dedicated userContextId container with strict OriginAttributes isolation for caches, storage, and connections. Password autofill and sensitive session data are disabled in the container when separation cannot be guaranteed, and POST requests trigger explicit user choice before reopening in the isolated context. The interface shows both a persistent container marker and a separate RU indicator only when Russian PKI is actively used, along with detailed security panels explaining the trust source. Updates, crash reporting, and automatic MAR mechanisms are intentionally omitted to avoid creating unverified trust chains for the distribution itself.
Survey of 254 Russian Domains Shows 89% DMARC Adoption but Highlights Gaps in Reporting and Subdomain Policies
A manual review of public DNS records across 254 prominent Russian domains from 17 sectors found strong baseline adoption of email authentication mechanisms. MX records appeared in 96.1% of domains, SPF in 93.7%, DMARC in 89.0%, and DKIM records via common selectors in 62.2%. Among domains with DMARC, 40.7% published a reject policy and 42.9% used quarantine, while 16.4% remained at none. Notably, 19% of DMARC-enabled domains lacked any rua address for aggregate reports, including 33 domains enforcing reject or quarantine. The study also identified cases of inconsistent policies between parent domains and subdomains, as well as SPF records ending in ~all paired with strict DMARC settings. Researchers emphasized that DNS data alone cannot confirm actual mail flow alignment or report consumption.
Server Outage Halts Vehicle Registration Across Smolensk Region
A technical failure on a unified server has temporarily suspended vehicle registration services in the Smolensk region of Russia. The outage affects the interdistrict traffic police department No.1 located on Lavochkina street, preventing new registrations from being processed. Regional UMVD officials confirmed that the problem impacts the single server used for the entire oblast's registration system. According to department head Maxim Zykov, the disruption is considered temporary, though no precise restoration timeline was provided. Applicants who submitted requests through the Gosuslugi portal will receive services in the first working days after the system is restored. The UMVD plans to issue an additional announcement once operations resume.