Ambient AI

Do ambient AI scribes need FDA regulation?

Two questions hide inside this one. Is an ambient scribe an FDA device today? — a statutory question with a reasonably clear answer. And should documentation tools face oversight given measured note-quality gaps? — an open policy question. This guide keeps them apart and maps each scribe capability to the statute it would or would not cross. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • Whether a scribe is an FDA device turns on the statutory device definition — software intended for the diagnosis, cure, mitigation, treatment, or prevention of disease. A tool that only transcribes and drafts a note for clinician review does not meet it.
  • The 21st Century Cures Act put two exits in the statute: software for administrative support of a health care facility, and qualifying clinical decision support where the clinician can independently review the basis for any recommendation. A pure scribe uses the first door.
  • So the legal answer today is fairly clear: a documentation-only ambient scribe sits outside FDA device regulation.
  • The open question is different: note quality is uneven — a blinded evaluation across 11 tools found human notes scored higher than AI notes on every case, and another detected hallucinations in 31% of AI notes — which is why some argue for oversight even where device rules do not reach.
  • The line moves as scribes add features. A tool that suggests diagnoses, orders, or codes drifts toward clinical decision support, and FDA's 2025 draft lifecycle guidance shows how it would be overseen if it crossed.

"Do ambient AI scribes need FDA regulation?" is really two questions wearing one coat. The first is legal and fairly settled: is a scribe an FDA-regulated device under the statute as it stands? The second is a policy judgment that is genuinely open: given that these tools produce clinically consequential documentation and can get it wrong, should they face some form of oversight at all? Conflating the two is how you end up with confident answers in both directions. This guide separates them, anchors the legal question to the statute and FDA guidance, and maps the policy tension to the evidence. As of July 2026.

This is general information, not legal or regulatory advice. Device classification is fact-specific and FDA guidance evolves. Confirm the regulatory status of any specific product with qualified regulatory counsel before relying on it.

What makes software a device

The whole question rests on one definition. Under the Federal Food, Drug, and Cosmetic Act, a device includes an instrument, apparatus, or software function "intended for use in the diagnosis of disease or other conditions, or in the cure, mitigation, treatment, or prevention of disease" 1. Intended use is the hinge. Software crosses into device territory because of what it is meant to do for a clinical decision — not because it is complex, not because it uses AI, and not because its output can be wrong. A tool whose intended use is to produce a draft note is doing something the definition was not written to capture. For the broader category this definition creates, see Software as a Medical Device.

The two doors out of the definition

The 21st Century Cures Act amended the statute to place certain software functions explicitly outside the device definition. Two of those exits matter here.

The first is administrative support. Software intended "for administrative support of a health care facility" — the statute lists functions such as the processing and maintenance of records — is excluded 2. Documentation is close to the paradigm case: a tool that captures the encounter and drafts the record is doing records work, not clinical decision-making.

The second door is the clinical decision support (CDS) carve-out, which FDA's 2022 final guidance interprets in detail 3. The statute removes qualifying CDS from the device definition when a set of criteria are met, and the decisive one for AI tools is that the software is intended "to enable the health care professional to independently review the basis for such recommendations," so the clinician does not rely primarily on the software to make the call 2. The guidance frames these as four criteria: the software does not analyze a medical image or a signal from a diagnostic device; it displays or analyzes medical information; it supports or provides recommendations to a clinician about prevention, diagnosis, or treatment; and it lets the clinician independently review the basis for those recommendations 3. Meet all four and the tool is Non-Device CDS; miss one — for example, by producing a directive the clinician cannot interrogate — and it can be a device. See clinical decision support system for how that boundary is drawn.

Where a pure scribe sits

Line up a documentation-only ambient scribe against those tests and the answer falls out cleanly. It does not analyze medical images or diagnostic-device signals; it processes speech into text. It does not, on its own, issue a diagnosis or a treatment recommendation; it drafts a record of what was said. And its entire output is a note the clinician reads, edits, and signs — the basis is fully reviewable because the clinician was in the room. On both the administrative-support ground and the CDS criteria, a scribe that only documents sits outside FDA device regulation. That is the reasonably clear legal answer today, and it is why no ambient scribe has been brought to market as a cleared device on the strength of its documentation function alone.

The open question the statute does not reach

Here is where the second question earns its place. Device regulation keys on intended use, so a tool can fall entirely outside it and still produce output that matters clinically and is sometimes wrong. The evidence that scribe output is uneven is now specific and peer-reviewed.

A blinded cross-sectional evaluation at the Veterans Health Administration compared 11 ambient scribe tools against human note-takers on five standardized primary-care cases, scored with a modified documentation-quality instrument. Across all five cases, "human-generated notes received higher overall modified PDQI-9 scores than AI-generated notes," with the largest gap on an acute low-back-pain case — human 43.8 versus AI 20.3 on a 50-point scale 6. A separate validated evaluation of an ambient scribe detected hallucinated content in 31% of AI notes versus 20% of clinician-written comparison notes 5, and a rapid review of the field catalogues the distinctive failure triad of "hallucinations, critical omissions, and misattribution" 7. Read AI hallucination in clinical contexts for why these errors are subtle enough to slip through a fast review.

None of that makes a scribe a device — but it is the substance of the argument that documentation tools should face some oversight, whether through quality standards, transparency requirements, or post-market monitoring, independent of the device question. A perspective on scaling these tools across varied settings makes the adjacent point: their "deployment in diverse care settings raises new challenges," and value depends on "thoughtful integration" and governance rather than on classification alone 4. The honest statement of the field is that the legal reach and the real risk are not the same size, and reasonable people disagree about how to close that gap.

It is worth being concrete about what oversight short of device regulation could mean, because "regulate them" and "leave them alone" are not the only options. A middle path borrows the machinery already used for deployed models without invoking the device framework: published note-quality benchmarks so buyers can compare tools on a common measure; disclosure of a tool's evaluated hallucination and omission rates; a requirement that clinician review remain in the workflow; and post-market surveillance of errors once a tool is in use. Much of that can be adopted voluntarily by health systems and professional bodies before any regulator acts, which is why the practical answer to "who watches these tools" is, for now, the organisations that deploy them. That places the burden on local governance — the subject of the argument in the scaling perspective 4 — rather than on a clearance process.

Where the line actually moves

The boundary is not fixed, because the product category is not standing still. The capabilities that would change a scribe's status are the ones that turn documentation into decision support:

Scribe capabilityToward or across the line?Governing test
Transcribe speech, draft the noteOutside — administrative/records functionDevice definition 1; admin support 2
Summarize history for clinician reviewOutside, if the basis stays reviewableCDS criteria 23
Suggest diagnoses or problemsToward — now "recommendations about diagnosis"CDS criterion 3 23
Draft orders or recommend codingToward — supports a clinical/billing decisionCDS criteria 3
Drive a decision the clinician cannot interrogateAcross — likely a deviceIndependent-review criterion 2

The determining factor is the fourth criterion: can the clinician independently review the basis for whatever the tool puts forward? A scribe that surfaces a suggested diagnosis with a transparent, checkable rationale may stay on the non-device side; one that issues a recommendation the clinician is expected to follow without being able to see why has crossed toward being a device.

A worked example makes the boundary tangible. Imagine one scribe that, alongside the draft note, lists "possible diagnoses to consider" with the specific phrases from the visit that prompted each, so the clinician can see exactly why the suggestion appeared and accept or dismiss it on the evidence. Now imagine a second tool that simply asserts a diagnosis and pre-fills orders to match, with no visible reasoning. Both "suggest a diagnosis," but they sit on opposite sides of the independent-review test: the first preserves the clinician's ability to interrogate the basis, while the second invites primary reliance on an opaque output. The feature label is the same; the regulatory status can differ entirely, which is why the design of how a recommendation is presented matters as much as whether one is made at all. For tools that do cross, FDA's January 2025 draft guidance, "Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations," lays out total-product-lifecycle expectations — transparency, bias management, and change control across the life of the product 8. The mechanism for managing ongoing model updates in that world is the predetermined change control plan, which lets a manufacturer pre-authorize defined changes rather than re-filing for each one. The coding-suggestion row above is also where this guide meets the billing risk covered in coding, billing, and upcoding risks.

How to read this

Three cautions. First, the legal answer is about intended use and marketing claims, so a vendor that adds decision-support features — or advertises them — can change a product's status even if the underlying model is similar; classification follows what the tool is held out to do. Second, "outside FDA device regulation" is not the same as "unregulated": privacy law, recording-consent law, and the standard of care all still apply, which is why deployment governance carries the weight the device framework does not. Third, guidance evolves — the CDS guidance and the AI lifecycle guidance are both live documents, and this page carries freshness triggers for the day either changes.

If you are deciding how to field one of these tools, the regulatory status is only the first checkpoint; the controls that actually protect patients live in the rollout. See the implementation checklist for those, the adoption statistics for how far the tools have spread, and ambient AI scribe for the underlying definition.

Sources and method

This guide is anchored to primary regulatory sources — the statutory device definition 1, the Cures Act software exclusions 2, FDA's 2022 CDS final guidance 3, and FDA's 2025 AI lifecycle draft guidance 8 — and to peer-reviewed evidence on note quality: a multi-tool quality evaluation 6, a validated hallucination study 5, a rapid review of failure modes 7, and a perspective on scaling 4. Every statutory phrase and figure is quoted from the source cited beside it, never from a summary. We revisit this page on a 180-day cycle and whenever a governing guidance changes. Nothing here is legal or regulatory advice.

Questions & answers

  • Does an ambient AI scribe need FDA clearance?

    As a general matter, a scribe that only listens, transcribes, and drafts a note for the clinician to review and sign does not meet the statutory definition of a device, so it falls outside FDA device clearance. The analysis changes if the tool starts making clinical recommendations. This is general information; confirm any regulatory question with counsel.

  • Why isn't a documentation tool that can hallucinate regulated?

    Because device regulation keys on intended use — diagnosis or treatment — rather than on whether output can be wrong. A tool intended only to draft documentation for clinician review sits outside the device definition even though its drafts can contain errors. That gap between real risk and regulatory reach is exactly the open policy question this page describes.

  • When would a scribe become an FDA device?

    When it stops merely documenting and starts supporting or providing recommendations about prevention, diagnosis, or treatment in a way the clinician cannot independently review — the criteria that define clinical decision support software. Features that suggest diagnoses, orders, or coding move a tool toward that line.

Sources

  1. 21 U.S.C. § 321(h) — Definition of 'device' under the Federal Food, Drug, and Cosmetic Act. Legal Information Institute, Cornell Law School. www.law.cornell.edu/uscode/text/21/321
  2. 21 U.S.C. § 360j(o) — Regulation of medical and certain decisions support software (FD&C Act § 520(o), as amended by the 21st Century Cures Act). Legal Information Institute, Cornell Law School. www.law.cornell.edu/uscode/text/21/360j
  3. U.S. Food and Drug Administration. Clinical Decision Support Software — Guidance for Industry and Food and Drug Administration Staff. September 2022 (Docket FDA-2017-D-6569). www.fda.gov/regulatory-information/search-fda-guidance-documents/clinical-decision-support-software
  4. Barriers and opportunities of scaling ambient AI scribes for clinical documentation across diverse healthcare settings. npj Digital Medicine. 2026. doi.org/10.1038/s41746-026-02554-0
  5. Assessing the quality of AI-generated clinical notes: validated evaluation of a large language model ambient scribe. Frontiers in Artificial Intelligence. 2025;8:1691499. doi.org/10.3389/frai.2025.1691499
  6. Rapid Evaluation of Artificial Intelligence Technology Used for Ambient Dictation in Primary Care: Comparing the Quality of Documentation of Artificial Intelligence-Generated and Human-Produced Clinical Notes. Annals of Internal Medicine. 2026. doi.org/10.7326/annals-25-02772
  7. Real-World Evidence Synthesis of Digital Scribes Using Ambient Listening and Generative Artificial Intelligence for Clinician Documentation Workflows: Rapid Review. JMIR AI. 2025;4:e76743. doi.org/10.2196/76743
  8. U.S. Food and Drug Administration. Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations; Draft Guidance. Federal Register, January 7, 2025. www.federalregister.gov/documents/2025/01/07/2024-31543/artificial-intelligence-enabled-device-software-functions-lifecycle-management-and-marketing