Ambient AI

Coding, billing, and upcoding risks of ambient AI notes

Ambient scribes write fuller notes — and a fuller note can support a higher bill. This guide connects the long regulatory record on documentation-driven upcoding to the specific new failure modes of AI-generated notes, and pairs each risk with a control that deployed programs actually run. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • Documentation-driven upcoding is a decades-old regulatory concern: Medicare E/M payments rose 48% from 2001 to 2010 as physicians billed higher-level codes across every service type, and in 2010 Medicare paid $6.7 billion — 21% of E/M payments — on claims that were incorrectly coded or unsupported.
  • Ambient AI notes are more thorough, which is the point — but a validated evaluation detected hallucinated content in 31% of AI notes versus 20% of clinician notes, so a fuller note is not automatically a truer one.
  • The core risk is that captured detail the encounter did not actually contain — hallucinated, cloned, or misattributed — ends up supporting a level of service that medical necessity does not.
  • Scribe adopters generated 1.81 more work RVUs per week than non-adopters in a study of 1.2 million encounters; that gain raises, rather than answers, whether the higher captured level is documented and warranted.
  • The control that deployed programs run is monitoring: auditing coding compliance against certified professional coders and keeping a clinician reviewing every note before it is signed.

Coding follows documentation. A clinician bills the level of service the note supports, and an ambient AI scribe changes the note — it produces a longer, more complete record of the encounter than most clinicians would write by hand. That is the feature. It is also the risk, because a fuller note can support a higher bill, and the question a compliance officer has to answer is whether the extra detail reflects what actually happened in the room. This guide sets the new technology against the long regulatory record on documentation-driven billing, names the specific risks, and pairs each with the control that deployed programs run. As of July 2026.

This is general information, not legal or billing-compliance advice. Coding rules, payer policies, and enforcement priorities change, and the right answer depends on your specialty, payer mix, and documentation. Confirm any coding or billing decision with your compliance office or counsel before acting.

The pattern is older than AI

Regulators have watched documentation drift upward for two decades, well before any model was in the room. The HHS Office of Inspector General found that Medicare payments for evaluation and management (E/M) services rose 48% from 2001 to 2010, from $22.7 billion to $33.5 billion, and that physicians "increased their billing of higher level codes for all types of E/M services" over that period 1. The same review identified roughly 1,700 physicians who consistently billed the highest-level codes — while carefully noting it "did not determine whether these physicians' claims were inappropriate," only that the pattern warranted attention 1.

A follow-on audit quantified the accuracy problem. For 2010, OIG estimated that Medicare "inappropriately paid $6.7 billion for claims for E/M services ... that were incorrectly coded and/or lacking documentation" — about 21% of E/M payments — with 42% of E/M claims coded incorrectly 2. The inpatient side shows the same shape: OIG found the number of Medicare stays billed at the highest severity level rose almost 20% from fiscal year 2014 to 2019, even as the average length of those stays fell, and flagged that the highest-severity stays "are vulnerable to inappropriate billing practices, such as upcoding — the practice of billing at a level that is higher than warranted" 3.

The through-line is that when documentation gets easier to inflate, billing tends to rise faster than the underlying care justifies. That is the lesson ambient AI inherits.

How a note becomes a bill

To see why a documentation tool is a billing tool, it helps to hold the mechanism in view. An evaluation-and-management level is not chosen at random; it is justified by what the note records. Under current rules a clinician selects the level based on the complexity of medical decision-making or on the total time spent on the encounter — and in either case the note is the evidence that the chosen level was warranted. A record that documents more problems addressed, more data reviewed, and more risk considered supports a higher level of medical decision-making; a record that logs more time supports a time-based level. This is why documentation and billing are the same conversation seen from two angles. Anything that changes what the note contains — a template, a macro, or now an ambient model — changes the raw material from which the bill is built, and puts the burden on the clinician to confirm that the captured complexity actually occurred.

What ambient AI changes

Two things are genuinely new, and they pull in opposite directions.

The first is thoroughness. Ambient notes capture more of the encounter than a rushed hand-written note, which can legitimately surface documentation that was always warranted but previously went unrecorded. In a study of 1,202,734 encounters, scribe adopters generated 1.81 more work RVUs per week than non-adopters 8 — a productivity signal that some of that captured detail is translating into billable work. On its own, that figure is neutral: it raises the question of whether the higher captured level is documented and warranted, and does not answer it.

The second is a new error profile. Older concerns centred on copy-paste and cloned notes; one review of that practice catalogued "regulatory concerns over the accuracy and medical necessity of billed services" and over fraud and abuse 4. Ambient AI adds a failure mode that dictation and templates did not have: fabrication. A validated evaluation of an ambient scribe detected hallucinated content in 31% of AI-generated notes versus 20% of clinician-written comparison notes 5, and a rapid review of the field describes the distinctive triad of "hallucinations, critical omissions, and misattribution" 6. A note can now be both more complete and less true — and both properties feed the coder. Read AI hallucination in clinical contexts for why these errors look plausible enough to survive a fast review.

The specific billing risks

The risks are concrete, and each has a mechanism worth naming.

RiskMechanismEvidence
Level inflationFuller documentation supports a higher E/M level than the visit's medical necessity warrantsHistorical E/M drift 1; improper-payment rate 2
Fabricated supportHallucinated exam findings or history appear in the note and back a billed element that did not occur31% vs 20% hallucination gap 5; failure-mode triad 6
Cloned detailContent carried across encounters makes each note look independently thoroughCopy-paste risk record 4
MisattributionFindings assigned to the wrong provider, visit, or problem distort what the note appears to supportFailure-mode triad 6
Severity creepCaptured detail pushes inpatient or problem severity upward without a matching change in careHighest-severity-stay trend 3

The common thread is that the note is the evidence. If an audit pulls the chart and the documentation contains detail the encounter did not produce, the defence that "the note supported the level" collapses — and the note was generated by a tool the clinician may have reviewed only lightly. That is the exposure.

It helps to picture the audit from the reviewer's side. A payer or a federal auditor does not see the visit; they see the chart, and they test whether the documentation supports the code that was billed and whether the service met medical necessity. A note that reads as unusually complete across a run of patients — the same rich review of systems, the same thorough examination, visit after visit — is exactly the pattern the OIG record describes as a marker worth scrutiny 1, and an ambient tool can produce that uniformity effortlessly. The danger is not that the tool sets out to inflate anything; it is that consistently fuller notes shift the statistical fingerprint of a practice upward, and some of that fullness may be detail the encounters did not actually contain. When the auditor asks the clinician to stand behind a specific documented finding, "the AI wrote it" is no defence — the signature on the note is.

The controls deployed programs run

The programs that take this seriously do not trust the output; they measure it.

Audit coding against real coders. One published deployment treated coding as something to verify continuously: it audited ICD-10 coding compliance "using an internally developed large language model" whose results were "assessed through correlation with certified professional coders," tracked on a real-time dashboard 7. The principle generalises — sample notes, compare the code the documentation supports against a qualified human coder's judgment, and watch the trend rather than a single snapshot.

Keep a clinician in the loop on every note. The single most important control is the oldest one: the clinician reads and edits the note before it is signed, and owns what it says. Because ambient notes carry the failure-mode triad above 6, review is where fabricated or misattributed detail is supposed to be caught — which only works if the review is real rather than a rubber stamp. See human-in-the-loop for what a meaningful review standard looks like.

Watch the productivity signal for what it means. An RVU gain like the 1.81 per week above 8 is worth tracking, but as a prompt for a question rather than a victory lap: is the additional captured level supported by the documentation and by medical necessity, on audit? A program that celebrates the revenue without checking the support is walking the path the OIG record describes.

Set an editing and attestation standard. The clinician review above only protects the practice if it is documented as a real step. A workable standard makes three things explicit: that the clinician read the full note, that the clinician corrected any content the encounter did not support, and that the clinician — rather than the tool — attests to the final record. Some groups add a light line noting that an AI tool assisted with drafting and that the clinician reviewed and edited the result; whether to include such language is a question for your compliance office, because it interacts with payer expectations and with the disclosure duties now appearing in some jurisdictions. What matters is that "reviewed and signed" means something specific and auditable, rather than a reflex click at the end of a busy clinic.

How to read this

Three cautions travel with everything here. First, none of the evidence shows that ambient scribes intend to upcode; the concern is that they make inflated or inaccurate documentation faster to produce, and the historical record shows where that leads 13. Second, the hallucination figures come from specific evaluations of specific tools 5, and error rates vary by vendor, specialty, and configuration, so treat them as a warning about the class of risk rather than a fixed rate. Third, the productivity and coding numbers describe associations in particular samples 8 and do not establish that any given claim was correct or incorrect — that determination is always specific to the chart and the payer's rules.

Because this is billing-adjacent, the operational takeaway is a governance one: before scaling an ambient scribe, decide who audits coding, how often, and against what human benchmark — and have your compliance office sign off on the review standard. For where that sits in a full rollout, see the implementation checklist; for how far these tools have already spread, see the adoption statistics; and for the tool itself, ambient AI scribe.

Sources and method

This guide pairs the federal audit record on documentation and E/M billing — three OIG reports on coding trends 1, improper payments 2, and inpatient severity 3 — with the peer-reviewed evidence on how ambient notes fail: a copy-paste risk review 4, a validated hallucination evaluation 5, and a rapid review of failure modes 6. The controls draw on a published monitoring playbook 7 and a large productivity study 8. Every figure is quoted from the primary report or paper cited beside it, never from a summary. We revisit this page on a 180-day cycle and whenever a new audit, study, or enforcement action lands. Nothing here is coding or legal advice.

Questions & answers

  • Do AI scribes cause upcoding?

    Not by themselves, but they change the raw material of a bill. Coding follows documentation, and ambient scribes produce fuller documentation — which can support a higher level of service. The risk is specific: when the extra detail is hallucinated, cloned, or misattributed, it can push a claim above what medical necessity supports. The evidence for that risk is the measured gap between AI-note and clinician-note error rates, not a claim that scribes intend to inflate coding.

  • Is a more thorough AI note a compliance advantage or a risk?

    Both, and which one depends on whether the detail is true. A thorough, accurate note supports the level billed and is easier to defend on audit. A thorough note padded with content the visit did not contain is the classic documentation-integrity problem that regulators have pursued for years, now produced faster. The determining factor is clinician review before the note is signed.

  • What control most reduces coding risk with ambient scribes?

    Two, together: a human-in-the-loop clinician reviewing and editing every note before signing, and an audit process that checks coding compliance against certified professional coders on an ongoing basis. Deployed programs that take coding seriously run both rather than assuming the output is right.

Sources

  1. U.S. Department of Health and Human Services, Office of Inspector General. Coding Trends of Medicare Evaluation and Management Services (OEI-04-10-00180). May 2012. oig.hhs.gov/reports/all/2012/coding-trends-of-medicare-evaluation-and-management-services/
  2. U.S. Department of Health and Human Services, Office of Inspector General. Improper Payments for Evaluation and Management Services Cost Medicare Billions in 2010 (OEI-04-10-00181). May 2014. oig.hhs.gov/reports/all/2014/improper-payments-for-evaluation-and-management-services-cost-medicare-billions-in-2010/
  3. U.S. Department of Health and Human Services, Office of Inspector General. Trend Toward More Expensive Inpatient Hospital Stays in Medicare Emerged Before COVID-19 and Warrants Further Scrutiny (OEI-02-18-00380). February 2021. oig.hhs.gov/oei/reports/OEI-02-18-00380.asp
  4. Weis JM, Levy PC. Copy, Paste, and Cloned Notes in Electronic Health Records: Prevalence, Benefits, Risks, and Best Practice Recommendations. CHEST. 2014;145(3):632-638. doi.org/10.1378/chest.13-0886
  5. Assessing the quality of AI-generated clinical notes: validated evaluation of a large language model ambient scribe. Frontiers in Artificial Intelligence. 2025;8:1691499. doi.org/10.3389/frai.2025.1691499
  6. Real-World Evidence Synthesis of Digital Scribes Using Ambient Listening and Generative Artificial Intelligence for Clinician Documentation Workflows: Rapid Review. JMIR AI. 2025;4:e76743. doi.org/10.2196/76743
  7. A Novel Playbook for Pragmatic Trial Operations to Monitor and Evaluate Ambient Artificial Intelligence in Clinical Practice. NEJM AI. 2025. doi.org/10.1056/AIdbp2401267
  8. Ambient Artificial Intelligence Scribes and Physician Financial Productivity. JAMA Network Open. 2026;9(1):e2553233. doi.org/10.1001/jamanetworkopen.2025.53233