AI-generated Replies Increase Physician Editing Workload

An ACL 2026 paper from Dartmouth researchers says AI-drafted patient replies can increase physician editing burden after evaluating 146,000 portal conversations from 10,105 patients. The team compared clinician responses with drafts from Claude, Gemini, ChatGPT, Llama, Aloe, and Qwen and found common misalignment: overly long answers, missing follow-up questions, and irrelevant or inaccurate medical details. News-Medical and Medical Xpress report that the work challenges a simple productivity story for clinical LLMs, because time saved on drafting can move into verification and correction. For healthcare AI teams, the practical takeaway is to instrument edit distance, missing-question rates, and clinician-specific adaptation before scaling patient-message assistants.
The Dartmouth study is useful because it measures the hidden cost of clinical LLM assistance: editing load. A draft that sounds fluent can still fail the workflow if physicians must spend more time correcting missing follow-up questions, irrelevant medical details, or patient-specific nuance than they would spend writing a reply themselves.
What happened
The ACL Anthology lists the 2026 paper, How Much Would a Clinician Edit This Draft? Evaluating LLM Alignment for Patient Message Response Drafting, by Parker Seegmiller, Joseph Gatto, Sarah E. Greer, Ganza Belise Isingizwe, Rohan Ray, Timothy E. Burdick, and Sarah Masud Preum. Medical Xpress, citing Dartmouth College, reports that the researchers analyzed 146,000 portal conversations from 10,105 patients in a large rural health system and evaluated drafts from Claude, Gemini, ChatGPT, Llama, Aloe, and Qwen.
Technical context
The paper frames the problem as alignment with individual clinician responses, not just generic answer quality. That distinction matters for deployment: patient portals require model output that asks the right clarifying questions, preserves clinical caution, and fits each clinician's communication style. According to Medical Xpress, the researchers found common draft problems including verbosity, missing follow-up questions, and irrelevant or inaccurate medical details.
For practitioners
Health systems should measure post-generation work before counting LLM replies as automation wins. Useful metrics include edit distance, omitted-question rates, unsafe or irrelevant detail rates, clinician-specific alignment, and time-to-send after review. Medical Xpress reports that adaptation to individual physicians improved accuracy by 33% and reduced editing by 26%, but the authors still emphasize clinician oversight.
What to watch
The next useful evidence is prospective workflow data: how much actual time clinicians spend editing drafts in live portal systems, how patients rate the final replies, and whether personalized adaptation holds up across specialties and patient populations.
Key Points
- 1The study shifts the question from whether LLM replies sound medical to whether they reduce clinician editing work.
- 2Missing follow-up questions and irrelevant medical details are workflow failures, even when a draft reads fluently.
- 3Healthcare AI teams should evaluate edit burden and clinician-specific alignment before expanding patient-message automation.
Scoring Rationale
This is a notable healthcare AI workflow study because it uses a large real portal-message corpus and measures editing burden rather than only answer quality. The impact is limited by the need for further prospective deployment data, but it is directly useful for clinical LLM adoption decisions.
Sources
Public references used for this report.
Practice with real Health & Insurance data
90 SQL & Python problems · 15 industry datasets
250 free problems · No credit card
See all Health & Insurance problems