Ambient AI scribes listen to a clinical encounter and draft the note. Four of them dominate the conversations clinicians and administrators are having right now, and the questions that follow are always the same: which one works, which one is safe, which one fits our systems. This page answers those questions the only honest way — as a capability matrix. It lays attributes side by side, ties every cell to a trial or to the vendor's own documentation, and declines to crown a winner. As of July 2026. For the underlying concept, see our glossary entry on the ambient AI scribe.
How to use this comparison
A ranking would tell you which tool is "best." A matrix tells you what each tool is documented to do and what independent evidence exists for it, and leaves the decision where it belongs — with the person who knows the setting. That distinction matters more here than in most software categories, because the evidence base is thin and lopsided: one product may have a randomized trial, and another only a product page. Treat an empty evidence cell as a prompt to ask the vendor, and read every number next to the study design that produced it — the discipline our guide on how to read an AI validation study lays out in full.
What the four tools are
Each of these is an ambient assistant that captures a conversation and produces a draft clinical note for a clinician to review and edit. They differ in lineage, integration, and how much independent study they have accumulated.
- Microsoft Dragon Copilot carries the lineage of the earlier Nuance DAX Copilot. Microsoft documents it as combining "conversational and ambient AI" with generative drafting, generating "draft clinical output from recordings for a clinician's review" 7.
- Nabla describes itself as an "ambient AI assistant" and documents support for 50+ specialties and multilingual capture 8.
- Abridge is the tool behind the largest published integrated-system deployment, the source of the ten-week and one-year records below 2310.
- Suki documents an "Ambient Clinical Intelligence" platform that spans documentation, dictation, coding support, and clinical question answering 9.
The capability matrix
Every filled cell below cites its source. A dash means the attribute was not stated in the sources cited here — read it as "ask the vendor," not as "absent."
| Attribute | Dragon Copilot | Nabla | Abridge | Suki |
|---|---|---|---|---|
| Vendor-independent randomized trial identified | Yes 1 | Yes 1 | — | — |
| Large real-world deployment study identified | — | — | Yes 23 | — |
| Design: ambient capture → clinician reviews draft | Yes 7 | Yes 8 | Yes 2 | Yes 9 |
| Documented EHR integration | Embedded via SDK or manual transfer 7 | Epic, Oracle Health, athenahealth, NextGen, Greenway 8 | Integrated-EHR deployment documented 23 | Epic, Oracle Health, athenahealth, MEDITECH 9 |
| Multilingual capture documented | Multi-party, multilingual 7 | Multilingual 8 | — | — |
| Scope beyond the note | Discrete orders, conditions, flowsheet data 7 | Documentation-focused 8 | Documentation-focused 10 | Dictation, coding support, Q&A 9 |
The most important row is the third: on the documented design, the four products converge. The tool drafts; a clinician reads, corrects, and signs. That shared human-in-the-loop step is what keeps a drafting error from silently becoming part of the record — and it is why the accuracy question below is about review burden, not autonomy.
What independent studies actually measured
Vendor documentation tells you what a tool is built to do. Independent studies tell you what happened when clinicians used it. Three bodies of evidence carry the most weight, and they do not all point the same way.
| Study | Design | What it measured | Headline result |
|---|---|---|---|
| Randomized trial 1 | 238 physicians, 14 specialties, DAX Copilot vs Nabla vs usual care, no vendor funding | Time-in-note (primary) | Nabla -9.5% vs control (p=.02); DAX Copilot no significant change |
| Integrated-system deployment 23 | Single large group, Abridge | Adoption and scale | 3,442 physicians / 303,266 encounters in 10 weeks; >2.5M uses in year one |
| Controlled time study 5 | ~1,800 clinicians | EHR and documentation time | ~13 min/day less in EHR; ~16 min/day less on documentation; 32% used it for >half of visits |
| Six-system burnout study 4 | Multi-site, pre/post | Self-reported burnout | Burnout fell from 51.9% to 38.8% after 30 days |
Read the randomized trial 1 carefully, because it is the only vendor-independent head-to-head here. On its primary outcome — time spent in the note — only Nabla achieved a statistically significant reduction (9.5%, p=.02); DAX Copilot did not differ significantly from control. On secondary, self-reported measures the two behaved alike: physicians using either scribe improved by 2.76 points on a standard burnout scale (p<.001). One trial at one academic system is a single data point rather than a verdict — but it is a pointed reminder that two tools marketed for the same job can produce different measured effects.
The deployment record 23 answers a different question: can a scribe run at scale? For one integrated group using Abridge, the answer was plainly yes — past 2.5 million uses in a year. That is the ceiling of a well-resourced rollout inside a single EHR, rather than the median any organization should expect. The ai scribe adoption statistics page holds the fuller set of deployment and time-savings figures.
The controlled time study 5 supplies the sober counterweight: savings are real but measured in a handful of minutes, and only about a third of users adopted the tool for the majority of their visits. The relief clinicians describe is genuine; the depth of use is uneven.
The average hides the variation
A single time-savings figure flattens a lopsided distribution. Because only 32% of clinicians used the scribe for more than half their visits 5, the average is dragged down by light users and up by heavy ones, and neither group resembles the mean. Setting matters as much as intensity: a specialty-specific evaluation among surgical residents found that this class of tool may help reduce documentation burden in that context 6, yet a result in one specialty and workflow rarely transfers cleanly to another. This is the practical argument for a local pilot over a borrowed benchmark — the question that decides value is whether a scribe saves time for your clinicians, in your specialties, at your adoption depth, rather than whether it saved time somewhere.
The data and consent layer the matrix does not show
Two attributes sit underneath every cell and deserve their own diligence. The first is recording consent: Microsoft's documentation is explicit that users should "obtain patient consent before recording the patient encounter" 7, and that requirement — with its jurisdictional variations — is a deployment obligation rather than a product feature. The second is where the audio and text go and who may process them; that governance question is invisible on a capability grid and decisive for compliance, so it belongs in any real evaluation alongside the measured effects above.
The accuracy question the matrix cannot settle
None of the deployment or time numbers matter if the draft is wrong. Ambient scribes carry a distinctive failure mode — content the clinician never said, and omissions of detail that was said — different from the transcription slips of older dictation tools. This is why every product in the matrix keeps a clinician reviewing each note, and why the review step is a feature rather than an inconvenience. The general shape of the problem is covered in our glossary entry on AI hallucination in clinical contexts. No published head-to-head has yet ranked these four on note accuracy under a common rubric, so the matrix leaves that row unfilled on purpose.
The regulatory frame
Ambient scribes generally sit outside the FDA's medical-device pathway, and the reason is structural. The agency's clinical decision support guidance turns in large part on whether a clinician can independently review the basis for the software's output rather than rely on it 11. Because a scribe produces a draft that a clinician reads and edits before signing, it typically falls on the non-device side of that line. That framing is not permanent: as products add ordering, coding, or recommendation features, the analysis can shift. Confirm the current status of any specific tool rather than assuming the category answer holds.
How to choose for your setting
Frame the decision as questions, not as a search for the highest-rated product.
- What is the primary problem? After-hours documentation, visit throughput, and clinician burnout are related but distinct; the studies above move each by different amounts.
- Which EHR do you run, and how deep is the integration? A tool that writes back into your EHR behaves differently from one that requires manual transfer 789.
- What languages do your patients and clinicians actually use? Multilingual capture is documented for some tools and unstated for others.
- Can you run a local, blinded pilot? The one randomized trial found divergent results for two popular tools — the strongest argument for measuring in your own clinics before committing.
- Who reviews the note, and how long does review take? The shared design puts a clinician in the loop; the real cost is review time, which a pilot can measure directly.
How to read this comparison
Four cautions travel with everything above. First, the evidence is lopsided — one product has a randomized trial, another only a product page — so absence of a study is not evidence of poor performance, and presence of one guarantees little. Second, the largest effects come from the best-resourced deployments, so selection is at work. Third, self-reported burnout and satisfaction gains lean on short windows and are consistently larger than the measured minutes. Fourth, features, integrations, and language support change on the vendors' timelines, so every cell here is a snapshot dated July 2026, and clearances and availability change — reconfirm before you rely on any single cell.
Sources and method
This comparison draws on one vendor-independent randomized trial 1, the ten-week and one-year deployment reports from a single integrated system 23, a large controlled time study 5, a multi-site burnout study 4, and a specialty-specific evaluation 6, alongside each vendor's own product documentation 78910 and the FDA's clinical decision support guidance 11. Every filled matrix cell is tied to one of these. We revisit this page on a 180-day cycle and whenever a new vendor-independent trial or a documented change in integration or language support lands. For the tools that answer clinical questions rather than write notes, see our clinical reference AI comparison.