Agentic AI

Documented agentic deployments in healthcare: a tracker

A running tally of the healthcare AI-agent systems that carry a published paper or the operator's own disclosure — what each one does, where the human sits, and what it actually reported. The field is loud; the documented set is small. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • Most 'agentic AI in healthcare' claims are benchmarks, pilots, or marketing — this tracker lists only systems with a published paper or the operator's own dated disclosure.
  • The best-documented deployment is a generative decision-support agent inside a Nairobi primary-care record, studied across 39,849 visits, with 16% fewer diagnostic and 13% fewer treatment errors among clinicians who had access.
  • The largest-scale disclosure is ambient documentation at an integrated Northern California group: over 2.5 million uses in year one, with a clinician editing and signing every note.
  • A real-hospital study connected an LLM to a live EHR through Model Context Protocol tools and reached near-perfect accuracy on simple infection-control retrieval, failing on complex time-dependent tasks.
  • One through-line holds across every documented row: a qualified human still decides. No fully autonomous clinical agent deployment is documented as of July 2026.

"Agentic AI" is the loudest phrase in health technology this year, and one of the least precise. Vendors describe autonomous agents that triage, order, schedule, and document; conference stages fill with roadmaps. Underneath the volume sits a much smaller question worth answering plainly: which agent systems are actually running in healthcare, and can you read the record yourself? This page tracks only the deployments that carry a published paper or the operator's own dated disclosure. Everything else — pilots without results, benchmarks, press releases without numbers — stays off the table until it earns a row. As of July 2026.

What earns a row

Three inclusion rules keep this tracker honest. First, the system has to touch real clinical work, rather than a simulated environment — that rule alone moves benchmarks into a separate section below. Second, there has to be a primary record: a peer-reviewed article, a preprint from the operator or its research partner, or the operator's own structured disclosure with dates and figures. A headline in the trade press does not qualify; the thing it summarizes might. Third, the entry has to describe something recognizably agentic in the working sense — a system that plans, retrieves, drafts, or acts through tools — or an adjacent workflow system whose documentation is strong enough to anchor the comparison. Naming a system here records that it is a documented subject of study. It is neither an endorsement nor a recommendation.

The tracker

Every figure below traces to the numbered source beside it. Systems are listed without ranking; the columns exist so you can judge fit for your own setting. Clearances, features, and availability change — treat each row as a dated snapshot, current as of July 2026.

System (operator)Setting & disclosureWhat it doesDocumentationOversight posture
AI Consult (Penda Health, with OpenAI)Nairobi primary-care network; study posted 2025Generative decision-support prompts inside the EHR, flagging possible documentation and decision errors during a visitPeer-reviewed preprint + operator disclosure 12Advisory only; clinician acts or dismisses — "preserving clinician autonomy" 1
Ambient AI scribe (integrated Northern California group)Outpatient care; year-one disclosure 2025Listens to the encounter and drafts the clinical noteOperator disclosure 34Clinician edits and signs every note
EHR-MCP retrieval agent (academic hospital)Infection-control workflow; study 2025LLM reads the live EHR through Model Context Protocol tools to answer clinical questionsPeer-reviewed preprint 5Supervised, retrospective evaluation; outputs checked
Polaris voice constellation (Hippocratic AI)Patient-facing phone conversations; report 2024Non-diagnostic voice agents for follow-up, education, and care coordinationVendor technical report 6Human clinicians supervise and escalate

AI Consult — the strongest deployment record

The clearest evidence that an agent-adjacent system changes care at scale comes from a primary-care network in Nairobi. A generative decision-support tool called AI Consult was built into the clinicians' electronic record, activating during a visit to flag potential documentation and decision-making errors. A real-world study compared care delivered by clinicians who did and did not have access across 39,849 visits at 15 clinics, and found 16% fewer diagnostic errors and 13% fewer treatment errors among those with the tool — while the system was designed as a safety net "activating only when needed and preserving clinician autonomy" 1. A companion mixed-methods evaluation recorded adoption climbing from about 4% to 47% of eligible episodes over eight months, with guidance surfaced through a tiered advisory interface a clinician could act on or wave past 2. This sits on the decision-support end of the agent spectrum — asynchronous, advisory, tool-triggered — and it is the best-documented example of the category touching real patients.

Ambient documentation — the largest-scale disclosure

The biggest numbers in the table are not from a diagnostic agent but from a documentation one. An integrated Northern California group disclosed that its ambient AI scribe passed over 2.5 million uses in its first year and entered a formal documentation quality-assurance program at that scale 3. The same group's earlier disclosure recorded 3,442 physicians using the tool across 303,266 encounters in the first ten weeks 4. Ambient scribing is a workflow-automation system rather than a planner that acts on the record, but its disclosure is the most detailed in healthcare, and the oversight model is unambiguous: the clinician edits and signs each note before it enters the chart. For the full evidence base on this category, see our AI scribe adoption statistics.

EHR-MCP — an agent reading a live record

The clearest demonstration that a tool-using agent can work against a real hospital's data comes from an infection-control study. Researchers connected an LLM (GPT-4.1) through a planning-and-acting agent to the live EHR using custom Model Context Protocol tools, then ran six tasks drawn from the infection-control team's real questions. The agent reached near-perfect accuracy on simple retrieval and degraded on a complex, time-dependent calculation, with most failures traced to wrong tool arguments or misread tool results; the authors frame the system as something that "may serve as a foundation for hospital AI agents" 5. It is a single-site, supervised evaluation rather than an at-scale rollout — which is exactly why it belongs here with that label attached.

Polaris — a disclosed patient-facing voice agent

A safety-focused voice-agent system for patient phone conversations was published as a technical report by its operator. Polaris is described as a constellation of a stateful primary agent plus several specialist support agents, evaluated by over 1,100 US-licensed nurses and over 130 physicians posing as patients, and reported to perform "on par with human nurses" on measures including medical safety and bedside manner 6. The evaluation used clinicians acting as patients rather than real-patient outcomes, and independent peer-reviewed outcome evidence remains limited — so the row records a detailed vendor disclosure of a deployed non-diagnostic system, read with that caveat.

Benchmarks are not deployments

Much of what circulates as proof of "agentic healthcare" is benchmark data, and it deserves respect in its own lane rather than promotion into this table. The reference EHR-agent benchmark runs 300 physician-written tasks in a FHIR-compliant virtual record; the best of 12 models completed 69.67% 7. A separate benchmark of clinical decision tasks measured the running cost of the agent pattern directly: more than ten times the token usage and more than twice the latency of the baseline models inside the agents 8. These numbers are informative about capability and cost. They are measured in simulation, and a simulated environment understates messy data, interruptions, and adversarial inputs. A benchmark score is a promise; a deployment is a result.

The oversight through-line

Read the table top to bottom and one pattern holds in every row: a qualified human still decides. AI Consult is advisory — a clinician acts on the flag or dismisses it 1. The ambient scribe drafts; a clinician signs 3. The EHR-MCP agent's outputs were checked in a supervised study 5. Polaris is non-diagnostic and escalates to clinicians 6. This is consistent with the wider literature: a 2026 scoping review of 43 included studies found published healthcare agents cluster into conversational, workflow-automation, and multimodal decision-support archetypes, overwhelmingly in pilots and research rather than at-scale production 9. The plain finding, then, is a double one. Documented agent deployments exist, and every documented one keeps a human in the loop. A fully autonomous clinical agent — one acting on consequential decisions with no clinician able to intervene — is undocumented as of July 2026.

Why so few systems make the table

Three structural forces keep this list short, and understanding them matters more than any single row. Disclosure is voluntary and asymmetric: an operator with a strong result has reasons to publish, while a stalled pilot rarely writes itself up, so the documented record is both sparse and skewed toward good news. Integration is expensive: reaching a live record safely, as the EHR-integrated agent evidence shows, takes standards, scoped access, and monitoring that most pilots never build past a demo. And the regulatory path for a multi-step, tool-using clinical agent is still forming — a deployment would be held to a stack assembled from general standards rather than one agent-specific rule, which our guide on agent safety frameworks maps in full. The net effect is a field where capability, measured in benchmarks, runs well ahead of documented deployment. A tracker that reported every vendor claim would be long and useless; one that reports only dated records is short and honest.

What would move a pilot onto this page

The bar for a new row is concrete, and worth stating so readers can apply it themselves. A system earns a place when three things exist together: a specific clinical setting rather than a lab; a primary record — a peer-reviewed article, an operator or research-partner preprint, or a structured disclosure with dates and figures; and enough methodological detail to tell what the number means. The most valuable future entries are the ones that close the gaps the current rows leave open: a deployment reporting hard patient outcomes rather than error-flag rates, an independent evaluation rather than an operator's own, and a system that documents how its tool-use failures — the wrong-argument mistakes seen in the real-hospital study 5 — are caught before they reach a patient. Each of those would strengthen the field's evidence base more than another capability benchmark.

How to read this tracker

Four cautions travel with the table. Selection is at work: the operators who publish detailed disclosures tend to be the best-resourced and most confident in their results, so the documented set skews toward success stories. Disclosure is not audit: an operator report and a preprint are primary records, and they are not the same as independent peer review or a regulator's finding. The categories blur: ambient scribing, advisory decision support, and tool-using retrieval are different animals, and lumping them under one word — "agent" — hides more than it reveals, which is why the "what it does" column matters more than the label. Finally, this is a snapshot. New rows will appear as papers and disclosures land; existing figures will age. We revisit this page on a 180-day cycle and whenever a new dated deployment record is published.

Sources and method

This tracker includes only systems with a primary record: two Penda Health / OpenAI reports on the AI Consult deployment 12, two operator disclosures of ambient documentation at scale 34, a real-hospital Model Context Protocol study 5, and a vendor technical report for a patient-facing voice constellation 6. Benchmark context comes from the reference EHR-agent benchmark 7 and an agent-cost benchmark 8, and the field's shape from a 2026 scoping review 9. Every figure is drawn from the source cited beside it, and each citation was checked to resolve before publication. For where these systems fit technically, see our guide on EHR-integrated agents; for the controls they run under, see agent safety frameworks for clinical settings.

Questions & answers

  • Are AI agents actually deployed in hospitals today?

    A handful are, and they are documented. As of July 2026 the systems with a published paper or the operator's own dated disclosure include a generative decision-support agent in a Kenyan primary-care record, ambient documentation at scale in a Northern California group, a real-hospital retrieval agent built on the Model Context Protocol, and a patient-facing voice-agent constellation. Most other "agent" claims are benchmarks, pilots, or marketing without a dated record behind them.

  • Is any clinical AI agent running without human oversight?

    No documented one. Every deployment in this tracker keeps a qualified clinician approving or able to override consequential actions — advisory prompts a clinician can dismiss, notes a clinician signs, retrieval a clinician checks. Full autonomy is a marketing claim, not a documented deployment, as of July 2026.

  • How is a "deployment" different from a benchmark score?

    A benchmark measures a system in a simulated environment; a deployment puts it in front of real clinical work. The reference EHR-agent benchmark had its best model complete about 70% of tasks in a virtual record — useful, and not the same as evidence from live patients. This tracker separates the two on purpose.

Sources

  1. Korom B, Amrose S, et al. AI-based Clinical Decision Support for Primary Care: A Real-World Study. arXiv:2507.16947. 2025. arxiv.org/abs/2507.16947
  2. A Mixed-Methods Evaluation of Clinician Experiences and Adoption Patterns of an EHR-integrated Generative AI-based Clinical Decision Support in Kenya. medRxiv 2025.08.14.25333740. 2025. doi.org/10.1101/2025.08.14.25333740
  3. The Permanente Medical Group. Ambient Artificial Intelligence Scribes: Learnings after 1 Year and over 2.5 Million Uses. NEJM Catalyst Innovations in Care Delivery. 2025. doi.org/10.1056/CAT.25.0040
  4. Tierney AA, Gayre G, Hoberman B, et al. Ambient Artificial Intelligence Scribes to Alleviate the Burden of Clinical Documentation. NEJM Catalyst Innovations in Care Delivery. 2024;5(3). doi.org/10.1056/CAT.23.0404
  5. Masayoshi K, et al. EHR-MCP: Real-world Evaluation of Clinical Information Retrieval by Large Language Models via Model Context Protocol. arXiv:2509.15957. 2025. arxiv.org/abs/2509.15957
  6. Mukherjee S, Gamble P, Ausin MS, et al. Polaris: A Safety-focused LLM Constellation Architecture for Healthcare. arXiv:2403.13313. 2024. arxiv.org/abs/2403.13313
  7. Jiang Y, Black KC, Geng G, et al. MedAgentBench: A Virtual EHR Environment to Benchmark Medical LLM Agents. NEJM AI. 2025. doi.org/10.1056/AIdbp2500144
  8. Liu Y, Carrero ZI, Jiang X, et al. Benchmarking large language model-based agent systems for clinical decision tasks. npj Digital Medicine. 2026. doi.org/10.1038/s41746-026-02443-6
  9. Njei B, Al-Ajlouni YA, Kanmounye US, et al. Artificial intelligence agents in healthcare research: A scoping review. PLOS One. 2026;21(2):e0342182. doi.org/10.1371/journal.pone.0342182