Tools that draft a reply to a patient's portal message — the AI text a clinician sees pre-written in the in-basket, ready to edit and send — spread through health systems quickly, and for once there is a real body of published evidence to read them by. This page sets that evidence side by side: what controlled studies and live deployments actually measured about empathy, readability, time, and clinician burden, and the one design choice every tool shares. It compares what the studies document. It names no best tool. As of July 2026.
What the studies measured
Four studies anchor the picture, and they are worth separating by design because each answers a different question.
The most-cited is a cross-sectional comparison that drew 195 exchanges from a public forum where a verified physician had answered a patient, then asked a licensed panel to compare those answers with a chatbot's. The panel preferred the chatbot's responses in 78.6% of evaluations and rated them higher on both quality and empathy 1. That result is striking and also limited: the questions came from a social forum rather than a clinic, and "preferred by a panel" measures perceived quality, which does not substitute for verified accuracy.
Two studies moved the question into the electronic health record. In a health- system evaluation of AI-generated in-basket drafts, usable AI responses were judged more empathetic than usable clinician responses (37.2% vs 16.5%) — and, in the same breath, less readable, a difference that matters most for patients with lower health or English literacy 2. A separate deployment measured the workflow effect: statistically significant reductions in clinician task load (61.31 to 47.26) and work exhaustion (1.95 to 1.62), but no change in reply, write, or read time 3. A fourth study of live use reinforces the design that all of them share — the AI produces a draft that a clinician reviews and edits before it is sent 4.
The comparison, attribute by attribute
The table reads the evidence across two tool types — a general chatbot answering patient questions, and AI draft generation built into the EHR in-basket. Each cell points to what a study documented, not to a vendor scoreboard. As of July 2026 — these are research findings that evolve; read each with its study design.
| Attribute | General LLM chatbot | EHR in-basket draft generation |
|---|---|---|
| Setting studied | Public forum questions, panel-rated 1 | Live patient portal messages in a health system 23 |
| Empathy vs clinician | Rated higher; preferred in 78.6% of evaluations 1 | Judged more empathetic (37.2% vs 16.5%) 2 |
| Readability | Not the primary measure | Lower than clinician drafts 2 |
| Time effect | Not measured in workflow | No change in reply, write, or read time 3 |
| Clinician burden | Not measured | Reduced task load and work exhaustion 3 |
| Human in the loop | Panel comparison, not deployed care 1 | Draft reviewed and edited before sending 4 |
| Safety evaluation | Thin across the field 5 | Thin across the field 5 |
Two patterns run across the columns. First, empathy is the consistent measured strength — every study that rated tone found the AI text warm, often warmer than the clinician's. Second, the operational payoff is softer than the empathy result implies: burden falls, but measured time does not, and readability can slip. A tool that makes a reply kinder and easier to face is worth something even if it saves no minutes; just do not buy it expecting the minutes.
The gap the studies do not fill: safety
The empathy findings are seductive, and they measure the wrong risk. Warmth is easy to rate; a missed red-flag symptom is not. The broadest look at how clinical LLMs are evaluated — a systematic review of 519 studies — found that only 5% used real patient-care data and that evaluation overwhelmingly optimized accuracy while calibration and uncertainty were assessed in barely 1.2% of studies 5. Patient-message drafting sits squarely in that blind spot: we know these tools sound empathetic; we have far less published evidence on how often they omit something important or state something wrong.
This is why the failure modes to watch are the quiet ones. Hallucination in clinical contexts — a confident, fluent sentence the record does not support — is more dangerous in a warm, readable message than in an obviously rough one, because fluency lowers the reader's guard. And a message channel that ingests patient text is a surface for prompt injection, where crafted input steers the model's output. Neither risk is measured by an empathy rating. WHO's guidance on large multi-modal models names the underlying hazard directly, warning that patient-facing generative outputs can generate misinformation and require human oversight 6.
That oversight is the load-bearing control. Every studied tool keeps a human in the loop: the clinician reads and edits the draft before a patient ever sees it. The empathy advantage is real; it earns the tool a place as a drafting aid, and the clinician's review is what keeps the draft from becoming an unverified clinical statement.
Where the benchmarks are heading
Because the deployment studies under-measure safety, the benchmark world has started to build the missing yardsticks. HealthBench grades models on 5,000 multi-turn conversations against 48,562 physician-written rubric criteria, explicitly scoring how a model behaves under uncertainty and in emergencies 7. MedHELM scores models across 121 clinician-defined tasks, patient communication among them 8. These are laboratory measures, and a high score still only licenses a local trial rather than clinical use — but they are moving evaluation toward the competencies an empathy rating skips.
How to choose for your setting
Questions to hold against any patient-communication drafting tool. The answers belong to your service line, your patient population, and your safety committee.
- What was actually measured — empathy, or accuracy? A demo that shows warmth has not shown safety. Ask for evaluation of omissions and errors, not tone alone 5.
- Will it save time, or only burden? The evidence so far separates these. Decide which one you are buying and measure it locally 3.
- Does readability fit your patients? If you serve lower-literacy or multilingual populations, test readability before rollout, since AI drafts have read lower on it 2.
- Is the human review real, and enforced? The control only works if the clinician genuinely reads and edits. Confirm the workflow makes editing the path of least resistance 4.
- What are the message-channel security controls? A tool that reads patient text needs defenses against prompt injection and data leakage 6.
For the adjacent story of how AI drafts the clinical note rather than the patient reply, see our AI scribe adoption statistics; for AI in teaching and assessment, see AI tools for medical education compared.
Sources and method
This comparison draws on four peer-reviewed studies of AI-drafted patient communication — a public-forum chatbot comparison 1, two EHR in-basket evaluations 23, and a live-use utility study 4 — set against a systematic review of how clinical LLMs are evaluated 5, WHO governance guidance 6, and two benchmarks building safety-oriented yardsticks 78. We present what each study measured, cite every attribute, and name no best tool; inclusion here is not an AIMOCS endorsement of any product. Study findings and benchmark scores are perishable; we revisit this page on a 180-day cycle and whenever a randomized trial or a safety-focused evaluation lands. As of July 2026.