The chief executive of the American Medical Association is publicly disputing a paper published this month in the association's own journal. The paper, by bioethicist Ezekiel Emanuel and venture investor Vinod Khosla, synthesizes hundreds of studies of AI in medicine and concludes that AI has already become an essential support tool for physicians and in many cases has begun to outperform them. Its sharpest claim is not about the future: on cognitive medical tasks, the authors write, treatment by AI alone is likely to be more effective than a physician alone or a physician working with AI. John Whyte, the AMA's CEO, is among the loudest critics of the work his organization's journal just ran.
Whyte's objections are methodological. Some of the studies gathered in the review were simulations rather than blinded experiments, the standard for medical research. Others contradict the authors' conclusions outright; he pointed to a paper in Nature from February that found most patients struggled to communicate with bots well enough to get the medical guidance they needed. Whyte allowed that the tools have potential, but said they should be used only inside a treatment plan directed by a physician.
Emanuel and Khosla tie their findings to how fast large language models have advanced since ChatGPT's public release in November 2022. Despite the technology's short history, they argue that AI already beats physicians at five basic tasks. Four of the five are eliciting the necessary medical information from a patient, reaching a diagnosis, running diagnostic workups, and managing chronic disease. The gap, they predict, will widen — AI is improving quickly while doctors shed some of their own skills by leaning on the same models. Emanuel told Wired the forecast runs four years out, and that it is hard to imagine AI will not be far ahead of a physician by then.
From there the authors reach their most contested position: that physicians should not intervene in the AI's work at all. The pairing of human and machine, the paper argues, actively degrades the result compared with the machine alone, and as AI improves, keeping a person in the loop will probably make care worse.
The deskilling half of that argument is not speculative. Experienced physicians already warn that medical students lean heavily on these tools, and the physicians are doing the same: internal data from OpenEvidence, an AI assistant built for clinicians, indicates roughly two-thirds of doctors in the United States use it. New clinical applications keep arriving — last week a team of neurosurgeons in the United Kingdom reported using an AI tool that assisted during an operation to remove a patient's brain tumor.
The load-bearing claim here is the one about the loop, and it is carrying more weight than the evidence under it can hold. "AI is useful" and "AI is often better than a doctor at a bounded task" are both defensible and neither changes anything structurally; hospitals will adopt the tools either way. "A doctor in the loop makes outcomes worse" is the claim that would rewrite liability, licensing and staffing, and it is precisely the claim most likely to rest on the simulation studies Whyte flags, because you cannot run a blinded trial of unsupervised AI care without first getting permission to deliver unsupervised AI care. The byline deserves the same scrutiny: a review arguing that the physician is the weak link, co-written by a venture investor, is not disqualified by that fact, but it is not a neutral reading of the literature either, and the AMA's journal published it without the association behind it.
Robert Wachter, who chairs medicine at the University of California, San Francisco, and is a friend of Emanuel's, thinks humans and AI currently do better together, though he grants that situations may arise in future where the human degrades the result. He also maintains physicians will always be indispensable, and calls the idea that AI will automate their work the "doorman fallacy" — the fear that automatic doors would make doormen obsolete. Doormen persisted because they take packages and provide a welcome a door cannot. The analogy is meant to reassure, and it does the opposite: it concedes that the core function was automated and defends what remains on the grounds of atmosphere. Nobody spends a decade in training to be atmosphere.
There is a recent precedent for how claims of this size end. Advocates of AI in education declared victory on the strength of a widely discussed study finding that ChatGPT could substantially improve learning outcomes. Springer Nature later retracted it over serious inconsistencies in the analysis. This paper is a synthesis rather than a single trial, so it will not fail that way — it will fail, or hold, one underlying study at a time, and the four-year clock Emanuel set is short enough that he will still be around to be held to it.