A provocative claim, not a settled clinical conclusion
The question of what remains for doctors as artificial intelligence improves has moved from speculative technology talk to a serious argument within medicine. A Perspective article published online in JAMA on August 17, 2026 argues that, for five central cognitive tasks, autonomous AI is likely to become better than both physicians working alone and physician-AI combinations. The tasks are taking a relevant history, forming a differential diagnosis, selecting tests, recommending guideline-concordant treatment and managing chronic disease.
The authors, led by University of Pennsylvania bioethicist Ezekiel Emanuel and including Curai Health chief executive Neal Khosla and investor Vinod Khosla, contend that the conventional aspiration of “AI-assisted physician care” may ultimately be surpassed by systems allowed to make the relevant cognitive decisions themselves. Their central proposition is counterintuitive: once an AI system is reliably better at a defined task, a clinician who can override it may introduce more error than benefit.
That is an important hypothesis, particularly because it challenges the reassuring assumption that a human in the loop is automatically safer. It is not, however, a demonstration that autonomous AI is ready to replace doctors in ordinary healthcare. The JAMA piece is a Perspective, not a prospective trial of autonomous medical practice, and its own authors acknowledge major limitations in the available evidence.
Strong benchmarks do not equal complete care
The argument rests on a rapidly expanding body of comparisons in which advanced models perform well against clinicians on narrowly specified diagnostic and treatment problems. Such results deserve attention. AI systems can retrieve and combine large volumes of medical information, apply guidelines consistently and examine possibilities outside an individual clinician’s immediate experience. In some settings, that may reduce missed alternatives, variation in routine decisions and delays in follow-up.
But benchmark success is not the same as managing an unstructured encounter. In a clinical vignette, the system is handed a curated history, selected test results and a well-defined question. In practice, the relevant facts may be incomplete, contradictory or never mentioned. A patient may not know which symptom matters, may have multiple conditions, may be unable to communicate clearly or may face practical barriers that change which treatment is appropriate.
A randomized Nature Medicine study published in February 2026 illustrates the gap. The researchers tested three general-purpose large language models with 1,298 members of the public handling medical scenarios. When models received the full scenario directly, they identified the condition accurately in almost 95 percent of cases on average. Yet people using those same models did not outperform a control group in choosing an appropriate course of action, and they identified relevant conditions less successfully. The information exchange between user and system was itself a point of failure.
This does not establish that physicians will necessarily outperform AI in every future conversation. It does show why claims based on isolated model performance must be distinguished from outcomes achieved by real people using a system in real circumstances.
The evidence for augmentation is becoming more practical
The alternative to both complacency and full autonomy is to study where AI meaningfully improves care teams. A pragmatic cluster-randomized trial published in Nature Medicine in 2026 tested a generative AI clinical decision-support system in routine primary-care workflows. It involved more than 9,000 analyzed patient encounters across 16 facilities.
The study did not find a statistically significant difference in its primary measure of treatment failure within 14 days. That is an essential restraint on exaggerated claims of transformation. Still, clinicians using the system produced more comprehensive notes and were more likely to document an appropriate diagnosis and treatment plan. The findings suggest that AI can improve elements of clinical process without yet proving that it delivers substantially better patient outcomes.
That distinction should shape adoption. Documentation, summarization, guideline retrieval, administrative communication and monitoring are not trivial parts of medicine; they consume time and can affect continuity of care. If AI reduces that burden safely, physicians may have more capacity for difficult diagnosis, communication, physical examination, procedures and coordination across fragmented services.
The most useful question is therefore not whether a model can beat a clinician on a test. It is whether a particular deployment improves outcomes, access, safety, fairness and patient experience compared with the care it replaces.
Why the physician role cannot be reduced to empathy
It is tempting to answer the automation challenge by assigning doctors the residual category of compassion. Empathy and difficult conversations are undoubtedly important, especially when patients confront uncertainty, serious illness or trade-offs between length and quality of life. But the professional role is wider than bedside manner.
Clinicians determine what information is missing, assess whether a result fits the patient in front of them, notice when the apparent problem is not the real problem and coordinate decisions involving families, specialists, carers and social services. They also hold responsibilities that are institutional as well as interpersonal: informed consent, advocacy, accountability, triage under scarcity and the challenge of unsafe or inequitable systems.
The American Medical Association and the Digital Medicine Society set out five enduring responsibilities for physicians in an August 2026 framework: preserving trust through human connection; exercising clinical judgment and accountability; leading changes in medical practice; stewarding the responsible use of technology; and advancing training and ethical decision-making. It is a professional vision rather than empirical proof, but it correctly identifies the parts of medicine that cannot be evaluated solely through a diagnostic accuracy score.
There is also a training problem. If early-career clinicians routinely accept machine-generated histories, diagnoses and plans without learning how to create and challenge them, they may lose the ability to identify when a system has failed. The JAMA authors themselves describe deskilling as a risk. Healthcare systems will need training models that treat AI as both a tool and an object of clinical scrutiny, rather than an unquestioned answer engine.
Regulation and accountability remain unresolved
Autonomous clinical AI would not only require stronger evidence; it would demand workable rules for responsibility. Who is accountable when an AI system makes a harmful recommendation, when its performance changes after an update, or when a hospital deploys it in a population unlike the one used for evaluation? The answer cannot simply be the clinician who was told not to intervene.
In the United States, the Food and Drug Administration’s final January 2026 guidance on clinical decision-support software clarifies which functions may be excluded from device regulation and which remain regulated device software. The framework reflects an important reality: intended use matters. A tool that supports a professional decision is not equivalent to one that directs a patient or delivers a clinical conclusion without meaningful review.
The FDA is also seeking input on regulatory approaches for generative-AI-enabled medical devices, including risk assessment, premarket evaluation and postmarket monitoring. That consultation underscores how unsettled the evidence and oversight questions remain. Reliable deployment will require evaluation beyond pre-launch testing, including monitoring for performance drift, cybersecurity problems, biased outputs and rare but severe failures.
A transition in work, not a verdict on doctors
The JAMA article is valuable because it forces medicine to take seriously the possibility that AI will eventually exceed human performance in defined cognitive tasks. Its prediction that autonomous systems could be deployable in some workflows by 2030 should be read as a contested forecast, not a clinical timetable.
Doctors are unlikely to retain authority merely because their role has tradition or social prestige. Where AI demonstrably improves a task, resisting it could harm patients and worsen already severe workload pressures. Conversely, removing clinicians simply because an AI excels in a simulated comparison would confuse technical capability with safe healthcare.
The likely transition will be uneven. Some repetitive, information-heavy functions may become substantially automated. Some specialties may change their training and staffing models faster than others. Yet the durable role of physicians will lie in setting the goals of care, validating whether the system is appropriate for a patient and context, taking responsibility for consequential decisions, and helping design institutions in which efficiency does not displace trust or equity.
The decisive standard is not whether there is anything left for doctors to do. It is whether new systems make care measurably safer, more accessible and more humane—and whether someone is clearly accountable when they do not.
Sources
- AI Has Human Doctors Asking: What’s Left for Us? — WIRED
- Will Autonomous AI Exceed AI-Aided Physicians as the Best Medical Care? — JAMA
- Reliability of LLMs as medical assistants for the general public: a randomized preregistered study — Nature Medicine
- Generative AI-enabled clinical decision support system in primary care: a pragmatic, cluster-randomized trial — Nature Medicine
- Considerations for the Regulation of Generative AI-Enabled Medical Devices — U.S. Food and Drug Administration



