Every ambient scribe vendor offers a return-on-investment calculator, and almost all of them share two flaws: they assume near-universal uptake, and they hide their formulas. This page does the opposite. It builds a return model from published figures, shows every input and every equation, and reports the result in the units the evidence actually measured — clinician-hours, visit capacity, and work RVUs — rather than currency. The monetary conversion depends on local wage and opportunity-cost figures that we deliberately leave to you. This is a model to test locally, never a promise. As of July 2026. For the underlying tool, see our glossary entry on the ambient AI scribe.
What this model is, and what it refuses to do
A return model is only as honest as its weakest assumption. So this one states its assumptions out loud and keeps the return in measured units. It expresses savings as time and capacity because those are what the controlled studies report; it declines to multiply by a currency figure because doing so would smuggle a local, highly variable number into a page that aims to be universal. Convert the final hours and RVUs to monetary value yourself, with your own finance team, and verify the result against your compliance and payer context before acting on it.
The model also refuses to assume full uptake. The single largest driver of a real-world return is how many clinicians actually use the tool, and how often — and the published adoption figures are sobering. Building those in is what separates a credible estimate from a brochure.
The published inputs
Every input below is a figure from a primary study, dated and sourced. Where a study reports a range, the range is shown.
| Input | Symbol | Published value | Source |
|---|---|---|---|
| Net EHR time saved per adopting clinician/day | — | ~13 min (3% relative) | 1 |
| Net documentation time saved per adopting clinician/day | M | ~16 min (10% relative) | 1 |
| Share of users adopting for >half of visits | — | 32% | 1 |
| Overall physician adoption where offered | a | ~45% | 2 |
| Additional patient visits per adopting clinician/week | V | ~0.5 | 1 |
| Additional work RVUs per adopting clinician/week | R | 1.81 | 2 |
| Per-encounter time-in-note reduction (tool-dependent) | — | 9.5% (one tool) / no significant change (another) | 4 |
| Burnout prevalence after 30 days | — | 51.9% → 38.8% | 3 |
Two inputs deserve a flag before any arithmetic. The minute savings in 1 are net — they are what remained after clinicians spent time reviewing and correcting each draft, so review time is already subtracted; adding it back as a separate cost would double-count. And the per-encounter reduction in the randomized trial 4 was tool-dependent: one scribe cut time-in-note by 9.5% while another showed no statistically significant change, so a per-encounter model must use the figure for the specific tool you pilot, never a category average.
The formulas
Three quantities, three equations. Let N be the number of clinicians offered the tool, a the adoption rate, so adopting clinicians A = N × a. Let D be clinical workdays per year and W clinical weeks per year.
- Clinician-hours saved per year = A × M × D ÷ 60
- Added visit capacity per year = A × V × W
- Added work RVUs per year = A × R × W
That is the whole model. Its transparency is the point: every term traces to a row in the table above, and you can substitute your own D, W, and — crucially — your own locally measured adoption rate the moment your pilot produces one.
A worked example
Take a 50-clinician group (N = 50) that offers the scribe to everyone. Using the published adoption rate (a = 0.45) gives about 22–23 adopting clinicians. Assume 220 clinical workdays (D) and 44 clinical weeks (W) per year, and the base documentation saving of M ≈ 15 minutes/day (mid-range of the 13–16 minute band).
| Quantity | Calculation | Result (per year) |
|---|---|---|
| Adopting clinicians | 50 × 0.45 | ≈ 22.5 |
| Clinician-hours saved | 22.5 × 15 × 220 ÷ 60 | ≈ 1,240 hours |
| Added visit capacity | 22.5 × 0.5 × 44 | ≈ 495 visits |
| Added work RVUs | 22.5 × 1.81 × 44 | ≈ 1,790 RVUs |
None of those three returns is interchangeable with the others: the hours are relief from after-visit work, the visits are throughput capacity that only materializes if the schedule is filled, and the RVUs are a productivity signal measured among adopters in one study 2. Report them separately, and resist summing them into a single hero number.
The two levers that move everything
A base case is a starting point rather than a forecast. The result swings most on two inputs, and a responsible model shows the swing rather than hiding it.
Adoption is the dominant lever. Because only 32% of users in the controlled study adopted the tool for more than half their visits 1, and overall uptake where offered ran near 45% 2, the number of genuinely active clinicians can be a fraction of the number licensed. Here is the same 50-clinician group across a plausible adoption-and-savings grid, in clinician-hours saved per year.
| Adoption rate → | 32% (16 adopters) | 45% (≈23 adopters) | 60% (30 adopters) |
|---|---|---|---|
| M = 13 min/day | ≈ 765 hr | ≈ 1,075 hr | ≈ 1,430 hr |
| M = 15 min/day | ≈ 880 hr | ≈ 1,240 hr | ≈ 1,650 hr |
| M = 16 min/day | ≈ 940 hr | ≈ 1,320 hr | ≈ 1,760 hr |
The top-left and bottom-right cells differ by more than a factor of two on the same 50 clinicians. Any calculator that reports a single confident figure without showing this spread promises certainty the evidence does not contain.
Setting is the second lever. A result in one specialty rarely transfers cleanly to another: a specialty-specific evaluation among surgical residents found this class of tool may help reduce documentation burden in that context 8, but the magnitude was particular to that workflow. Your own specialty mix, note complexity, and baseline documentation load all move M.
The return the model does not monetize: burnout
Some of the most consistent evidence is the hardest to put in a formula. Across six health systems, the share of clinicians reporting burnout fell from 51.9% to 38.8% after thirty days with an ambient scribe — roughly 74% lower odds 3. That is a real return on retention, recruitment, and care quality, and it is measured, but it comes from short-window self-report and resists conversion into hours or RVUs. Keep it beside the capacity figures as a distinct line, and weigh it with the same skepticism as the rest — the discipline our guide on how to read an AI validation study sets out.
Feasibility: can the tool run at the scale you are modeling?
A capacity model assumes the tool can actually carry the volume. The published ceiling is reassuring on that narrow point: one integrated group put a scribe across 303,266 encounters in ten weeks 5 and past 2.5 million uses in its first year 6. That establishes scale is achievable inside a single EHR with a coordinated rollout — while remaining, as ever, the best-resourced case rather than the median. The AI scribe adoption statistics page holds the fuller deployment record, and the implementation checklist for medical groups covers the rollout mechanics this model assumes are in place.
A second lens: the per-encounter model
The daily-net model above is the most defensible, because its inputs were measured as daily totals that already net out review time. But leaders often want a per-encounter view, and the randomized trial supplies one — with a sharp caveat. In that trial, one scribe cut time-in-note by 9.5% while another showed no statistically significant change 4. The per-encounter equation is:
- Time-in-note saved per year = A × E × t × r
where E is encounters per clinician per year, t is baseline minutes in the note per encounter, and r is the tool-specific fractional reduction. The load-bearing term is r: applying the 9.5% figure to a tool that has never been measured — or, worse, a category average — converts the model into a guess. Only a local pilot, or an independent trial of the specific tool, yields an r you can trust. And because the reduction is on time-in-note rather than on the whole encounter, this lens captures a narrower slice of the day than the daily-net model, so treat the two as bounds rather than adding them together.
What this model deliberately leaves out
A transparent model is honest about its own edges. Three real costs sit outside the equations above and belong in any full appraisal:
- Ramp time. The measured savings appear after clinicians have learned the tool; the first weeks can run net-negative as the review habit forms.
- Light users can lose time. Because only 32% of users adopt for the majority of visits 1, an occasional user may spend longer reviewing an unfamiliar draft than they save — a group the average conceals.
- Implementation and oversight effort. Rollout, EHR integration, consent workflows, and the documentation quality-assurance program a large deployment eventually needs 6 all consume staff time this model does not count. The implementation checklist for medical groups covers that side.
How to read this model
Four cautions travel with every figure above. First, it is a model: it propagates published averages through arithmetic, and averages hide the lopsided distribution of light and heavy users. Second, the inputs come from specific populations — large, often well-resourced groups — so your baseline documentation load, specialty mix, and EHR determine whether M lands at 13, 16, or elsewhere. Third, the returns are in different currencies of value — hours, visits, RVUs, burnout — and collapsing them into one number destroys the information that makes each decision-relevant. Fourth, the human review step is load-bearing and already priced into the minute savings; the human-in-the-loop design is what keeps ambient-note failure modes — hallucination and omission 7 — out of the record, and its time cost is a feature of the net figure. Because this page touches financial and productivity decisions, run the monetary conversion and the payer-context check with your own finance and compliance functions before relying on any output.
Sources and method
This model is built from a controlled time-and-visit study 1, a work-RVU productivity study 2, a six-system burnout study 3, the one vendor-independent randomized trial 4, the ten-week and one-year deployment records from a single integrated system 56, a real-world evidence synthesis of note failure modes 7, and a specialty-specific evaluation 8. Every input in the table traces to one of these; every formula is shown in full so the model can be re-run with local numbers. We revisit this page on a 180-day cycle and whenever a new controlled study revises an effect size. For an attribute-by-attribute view of the tools themselves, see our AI scribe head-to-head comparison. As of July 2026.