A randomized clinical trial published in NEJM AI enrolled 44 physicians in Pakistan who had completed a structured AI literacy programme. Control-group physicians received accurate diagnostic suggestions from GPT-4o; treatment-group physicians received suggestions containing deliberately introduced errors. Treatment physicians scored 73.3% on diagnostic reasoning tests versus 84.9% in the control group. Top-choice diagnosis accuracy fell from 90.5% (control) to 76.1% (treatment) — a decline of 14.4 percentage points. A companion study from MIT, published August 4, 2026, found that benefits from LLM diagnostic assistance were significantly differentiated by user expertise: clinicians were more likely to catch AI errors than non-experts, but neither group was reliably immune. The structural finding: AI literacy training reduces but does not eliminate automation bias. When the LLM sounds confident and is wrong, trained physicians still follow it at a rate that is clinically significant.
1. What the Trial Found
The trial, published in NEJM AI under the title “Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy,” enrolled 44 physicians based in Pakistan who had completed a structured AI literacy training programme. [Established — NEJM AI, “Automation Bias in Large Language Model–Assisted Diagnostic Reasoning among Physicians Trained in AI Literacy — A Randomized Clinical Trial,” NEJM AI doi:10.1056/AIoa2501001. Publication date: 2025–2026.] The design was a randomized controlled trial with two arms: the control arm received correct diagnostic suggestions from ChatGPT 4o; the treatment arm received suggestions containing deliberately introduced diagnostic errors.
The results were statistically significant. Physicians in the treatment group — those given the wrong AI advice — scored 73.3% on a structured diagnostic reasoning assessment, compared with 84.9% in the control group. [Established — NEJM AI, ibid.; also reported by Belgian Journal of Hospital Medicine (BJH), “Automation Bias Persists in AI-Trained Physicians,” August 2026.] The top-choice diagnosis accuracy measure — whether the physician’s first-choice diagnosis was correct — declined from 90.5% in the control group to 76.1% in the treatment group: a 14.4 percentage point difference attributable to automation bias.
Automation bias is the documented tendency of humans to defer to automated system outputs over their own judgement, particularly when the system presents its recommendation with apparent confidence. The finding is not new in kind — automation bias has been documented in aviation, nuclear power operation, and radiology for decades. What is new is a rigorously designed randomized trial demonstrating its presence and magnitude in AI-assisted clinical diagnosis, specifically among physicians who had undergone the kind of structured AI training that is now being deployed across health systems globally as the solution to AI safety concerns.
2. Why AI Literacy Isn’t Enough
The trial’s most important finding is not that physicians fell for the wrong AI advice. It is that physicians who had been specifically trained to be sceptical of AI advice still fell for it at a rate that translated to a 14-point accuracy gap. The literacy programme was not a cosmetic intervention; it was a structured curriculum designed to improve AI-critical evaluation. It helped — the treatment-group performance was still within a plausible range for diagnostic medicine — but it did not close the automation bias gap.
The mechanism is cognitive, not motivational. A physician who intends to critically evaluate AI suggestions still faces the problem that the AI’s output occupies attention and working memory in a way that reconfigures how subsequent evidence is processed. When the LLM’s suggestion is confident and plausible-sounding, the physician’s cognitive resources are partially committed to evaluating whether to accept it rather than to generating an independent differential diagnosis. The training addresses the physician’s attitude toward AI; it does not address the cognitive load architecture that makes the bias persistent. [Assessed with moderate confidence — mechanistic interpretation of the trial findings; the specific cognitive mechanism is inferred from the automation bias literature, not directly measured in this study.]
3. The MIT Study: The Non-Expert Problem Is Larger
A companion finding from MIT, published August 4, 2026, examined how AI diagnostic assistance affected users differentiated by clinical expertise. [Established — MIT News, “Benefits of medical AI assistance vary based on user expertise,” 4 August 2026.] The study found that non-clinicians — the patients, administrators, and non-specialist users who increasingly encounter AI diagnostic tools in direct-to-consumer and telehealth settings — deferred to LLM diagnostic suggestions even when they held doubts about the recommendation. Clinicians were more likely to catch AI errors, but both groups showed susceptibility.
The practical significance is this: the healthcare AI deployment landscape is not limited to physicians using clinical decision support tools in hospital settings. It increasingly includes AI-powered symptom checkers, direct-to-consumer diagnostic aids, and telehealth triage tools used by people without clinical training. The MIT finding suggests that those users — who face no literacy training, no structured AI education, and often no pathway to escalate an AI recommendation to expert review — face a substantially larger automation bias risk than the physician population in the NEJM AI trial. [Established — MIT News, 4 August 2026; analytical inference from the comparison of the two study populations.]
4. What Structural Redesign Means
The trial’s authors recommend pairing LLM use with “workflow safeguards, verification steps, and human oversight.” [Established — NEJM AI trial; BJH reporting.] These are correct in direction but require unpacking. The failure mode the trial documented is not that physicians lacked oversight capacity; it is that the LLM’s confident presentation of a wrong answer restructured how the physician used their oversight capacity. Adding a verification step that the physician performs after receiving the LLM output faces the same problem as the literacy training: it operates downstream of the cognitive reconfiguration the LLM’s output has already caused.
The structural redesign implied by the evidence has a different character. It involves changes to when and how the LLM output is surfaced in the clinical workflow — not just adding a check at the end of the existing process. Proposals in the automation bias literature that may be relevant include: requiring the physician to complete an independent differential diagnosis before the LLM output is displayed; presenting the LLM’s reasoning rather than its conclusion first, forcing engagement with evidence before the verdict; and structuring the interface to display explicit uncertainty ranges rather than a single confident recommendation. None of these is standard practice in currently deployed clinical AI tools. [Assessed with moderate confidence — based on automation bias mitigation literature; the specific effectiveness of these interventions in clinical AI settings has not been validated in RCTs of equivalent quality to the NEJM AI study.]
5. The Broader Implication
The healthcare industry has invested heavily in two arguments for AI safety: that AI systems are improving rapidly enough that their errors are becoming rare, and that human oversight provides a reliable check when they are not. The NEJM AI trial does not challenge the first argument. It challenges the second. If a physician who has been trained in AI literacy still follows the AI when it is wrong, the human-oversight backstop is weaker than the deployment decision assumed.
This matters beyond medicine. The same automation bias dynamic — confident AI output reconfiguring human cognitive processing — applies in any high-stakes human-AI collaboration where the human is expected to serve as a final check: legal document review, financial fraud detection, air-traffic control, cyber-threat analysis. The medical trial is unusually clean because it used a randomized design and a measurable outcome (diagnostic accuracy). But the mechanism it documents is not unique to medicine. [Assessed — cross-domain inference; the automation bias literature is robust across professional domains, though no equivalent RCT exists in those specific fields.]
The appropriate response is not to pause clinical AI deployment. The MIT study found that AI assistance does produce accuracy gains for clinicians — the gains are just unevenly distributed and partially offset by bias when the AI is wrong. The appropriate response is to treat the NEJM AI finding as design specification: any AI system deployed in a high-stakes human-oversight environment needs to be engineered around the documented failure mode of automation bias, not around the assumption that trained humans will reliably catch AI errors when they occur.
Bottom line: AI literacy training for physicians is necessary. This trial demonstrates that it is not sufficient. A 14-percentage-point drop in diagnostic accuracy when AI advice is wrong — in a population that has been specifically trained to be critical — is a design constraint, not a training failure. The constraint requires structural solutions at the workflow and interface level, not further investment in the same class of literacy intervention that the trial shows cannot close the automation bias gap on its own.