In a randomized trial published in JAMA Network Open in October 2024, fifty American physicians were tested on complex diagnostic cases: GPT-4 working alone scored a median 92%, physicians with GPT-4 available to them scored 76%, and physicians using conventional resources scored 74%.
The model beat the doctor. The model also beat the model-plus-doctor.
AI Key Takeaways
A randomized trial showed GPT-4 working alone outperforming both physicians with GPT-4 (92% vs 76%) and physicians without it — the model beat the doctor, and the model beat the model-plus-doctor.
Non-maleficence, medicine’s oldest ethical commitment, has a second edge: a duty not to leave available accuracy on the table.
The standard of care is not set by regulators; it is set retrospectively by what a reasonable practitioner would have done — and one case is enough to move it.
A follow-up in Nature Medicine found improvement on management reasoning, which complicates the picture and does not change the core: the performance gap is widening, not closing.
The word for this is non-maleficence
Medical ethics rests on four principles, and one of them is non-maleficence, the duty not to inflict harm. It is the oldest commitment in the profession, older than consent, older than the codes, and it is what the phrase primum non nocere refers to. In law the same idea appears as the duty of care, and a breach of it is what negligence means.
Non-maleficence is normally invoked to restrain intervention. Do not operate unnecessarily. Do not prescribe what carries more risk than benefit. Do not experiment on a patient for your own curiosity.
A duty not to harm is also a duty not to leave available accuracy on the table. If a documented, cheap, widely accessible tool would have caught what you missed, the harm to your patient was avoidable, and you are the reason it was not avoided.
Is it defensible for a physician to decline a tool that outperforms them, when the consequence of declining falls on the patient rather than on them?
Looking at the numbers, and numbers are all that matters (to me) when they are not just numbers, but people, I believe it is not. And I think the profession will discover this through litigation rather than through debate. The sequence is easy to predict diagnosis is missed. The family’s lawyer runs the case through a model, which produces the correct answer in ninety seconds. He then produces the trial data showing that this was foreseeable at the population level, and asks the physician on the stand why he did not use it.
There is no good answer to that question. “I preferred my own judgment” describes the harm rather than excusing it.
Once one case lands, the standard of care moves. Standards of care are not set by regulators; they are set by what a reasonable practitioner would have done, and what a reasonable practitioner would have done is established retrospectively, in court, by reference to what was available. Nothing needs to be legislated. Nothing needs to be adopted. The obligation arrives through liability, and it arrives whether or not anyone in the profession wants it.
The endpoint is a reversal of the current position. Today, AI in medicine is discussed as a tool a physician may choose to use, with the physician as the authority and the machine as the assistant. The reversal makes the machine the baseline and the physician the exception, required to document why he departed from it. Prescribing without that check becomes what prescribing while refusing to read the chart is now.

AI diagnostics outperform human doctors every time, yet patients hesitate to rely on them. This hesitation is grounded in a rational understanding of accountability. We trust doctors not because they are infallible, but because of the accountability infrastructure surrounding them: board certifications, malpractice liability, and revocable licenses.
The mechanism in the 2024 trial was not straightforward machine superiority. The physicians who had GPT-4 largely did not revise their view when it disagreed with them. They anchored on their own initial impression, which is why the assisted group barely outperformed the unassisted one. This is a finding about human deference, and it cuts in an uncomfortable direction: it suggests that giving doctors the tool is insufficient, and that capturing the accuracy would require making the machine primary and the human the reviewer. That is a larger change than adding software to a workflow.
A follow-up in Nature Medicine found the opposite result on a different task. Ninety-two physicians given model access alongside conventional resources did improve the quality of their reasoning on patient management. Diagnosis and management come apart here.
And this gap is not closing. When OpenAI’s o1-preview was run against the same six cases, it scored a median 97%, against GPT-4’s 92% (whatever the correct account of why doctors fail to use these systems well), the distance between unaided human performance and machine performance is widening, and the duty not to harm is indifferent to the explanation.
Soon it will be malpractice to diagnose or prescribe without AI oversight.
NA: AI-assisted tools were used for transcription, reference formatting, and language editing. All intellectual content and conclusions remain solely the author’s.







I love when AI and ethics intersect. Great piece!