Agentic AI

Agents in nursing workflows: documentation, handover, and the virtual-nursing evidence

Two very different things get marketed as the "AI nurse": documentation assistants that reformat what a nurse says, and virtual-nursing programs that move tasks to a remote colleague. The measured time savings, the completeness gains, and why one of these has much stronger evidence than the other. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • The burden is documented and large: in one 2025 study nurses logged 631 to 875 flowsheet entries per 12-hour shift — about one a minute — and spent 31% of the shift in flowsheets.
  • Documentation agents show real, replicated time savings: an LLM constrained to reformat nurse input into a structured handover cut per-patient documentation time from 3.4-4.3 to 1.2-2.5 minutes across three hospitals.
  • That same team scope-limited the agent deliberately — it stopped letting the model structure vital signs because of hallucination risk, and required a nurse to review and confirm every note.
  • The other 'AI nurse' is virtual nursing, where a 2025 study found higher documentation completeness (80% vs 59% on admissions) — but the design was observational and confounded, so the gain cannot be pinned on the technology.
  • The professional standard is explicit: the American Nurses Association holds that AI augments rather than replaces nursing care, and that biased training data carries its bias into practice.

Walk a hospital ward and the nurse in front of you is, for a large share of the shift, looking at a screen. Nursing has become one of the most documentation-heavy jobs in healthcare, and that burden is the opening any AI vendor pitches into. The pitch comes in two very different shapes, though, and conflating them is the most common mistake buyers make. One is a documentation assistant that helps a nurse chart faster. The other is a virtual-nursing model that moves whole tasks to a remote colleague. They have different evidence, different risks, and different things to prove. This guide separates them and holds each to what the published record supports. As of July 2026.

The burden these agents are aimed at

The problem is well measured, and the numbers are striking. A 2025 study at a large academic health system found that nurses complete between 631 and 875 flowsheet entries per 12-hour shift — close to one entry every minute for the entire shift — and that electronic-record vendor data showed nurses spending 31% of a 12-hour shift working in flowsheets 1. The nurses in that study named "redundant and nonmeaningful documentation" as a primary frustration 1: a meaningful share of the burden is duplicated data entry, not clinical thinking.

A 2024 mixed-methods study of acute and critical-care nurses put a time figure on it, measuring documentation burden of 78.5 to 88.4 mean minutes across five units and identifying poor electronic-record usability and data redundancy as a major contributing factor, with flowsheets rated the most burdensome component of all 2. That study also ran a standard usability instrument, and the results point directly at where an agent should and should not go: only the medication and free-text note components scored above the accepted usability threshold, while flowsheets and care plans fell well below it 2. The most burdensome, least usable component — the flowsheet — is also the one built from discrete, structured values, which is exactly the kind of data a language model is most likely to get subtly wrong. The authors add a caution that shapes everything downstream: usability interventions "must be nurse-endorsed" 2. An agent imposed on a workflow the nurses do not trust reproduces the original problem in a new form. This is the same documentation-load story that ambient tools tell on the physician side, tracked in our AI scribe adoption statistics.

What a documentation agent measurably does

Here the evidence is the strongest in this whole area, and it is worth being precise about why. A 2026 study across three hospitals in Taiwan integrated a large language model into nursing handover documentation and measured the before-and-after in routine use. The tool was constrained to reformat nurse-provided input into a structured ISBAR handover note — Identify, Situation, Background, Assessment, Recommendation. Per-patient handover documentation time fell from 3.45-4.32 minutes to 1.17-2.54 minutes, a 26-73% relative decrease, translating to an estimated 474-981 nursing hours saved each month across the three sites 3.

The design choices behind that result matter more than the headline. This was an ambient-adjacent documentation aid deliberately kept on a short leash. The team discontinued LLM structuring of vital signs after judging the hallucination risk unacceptable, and they required nurse review, correction, and confirmation before any note was finalized 3. In other words, the agent was scoped to a task where reformatting is safe, pulled off a task where fabrication is dangerous, and never allowed to be the final author. The time saving is real because the design was conservative — a worked example of the human in the loop as a feature rather than an afterthought. That is the template to copy: constrain the agent to reformatting what a clinician supplies, keep it away from inventing clinical values, and make a nurse the last set of eyes.

It is worth naming why the handover was a good first target and the flowsheet was not. A handover note is narrative synthesis — the nurse already holds the facts and needs them organized, so the model's job is compression and structure, where a mistake is visible and recoverable. A vital sign or a medication value is a discrete datum with a right answer, where a plausible-but-wrong number can slip past review and into a decision. The general lesson generalizes past nursing: agents are safest on tasks that reorganize information a human already verified, and most dangerous on tasks that originate a clinical value. Ambient documentation for nurses — a microphone capturing a bedside interaction and drafting the note — sits between the two, and it inherits the same rule: the physician-side scribe evidence in our AI scribe adoption statistics shows measurable time savings paired with a persistent tail of fabricated and omitted detail, which is why every credible deployment keeps a clinician confirming the draft.

The other "AI nurse": virtual nursing

The second shape is different in kind. Virtual — or tele — nursing puts an experienced nurse on a screen, handling admissions, discharges, education, and documentation for the bedside team, and it is increasingly described in AI terms even when the core is a remote human. A 2025 observational study of 111 encounters (81 virtual, 30 in-person) found that virtual nurses achieved higher documentation completeness than bedside nurses: 80.2% versus 58.8% on admissions, with large gaps on specific tasks — home medication review at 78.8% versus 25%, and discharge education at 79.2% versus 27.2% 4.

Those are meaningful numbers, and they are exactly the kind that need reading carefully. The study was observational and cross-sectional, and the virtual nurses were dedicated to admission and discharge work while the bedside nurses were doing everything at once 4. When one group is protected to do a task and another is interrupted through it, the protected group will score higher on that task regardless of the technology. The completeness gain is a genuine finding about a role design — giving a nurse the space to finish admission paperwork — that happens to be delivered over video. It cannot be read as evidence that an AI agent produced the improvement, because no controlled comparison isolates the technology from the reallocation of attention. That distinction is the entire subject of our guide on how to read an AI validation study.

The standard both are measured against

Nursing has a professional position on all of this, and it predates the current wave. The American Nurses Association's position statement, adopted in December 2022, states that "AI does not replace good nursing care" and that "AI augments, supports, and streamlines expert clinical practice" 5. It also draws a specific warning about bias: data "mined from domains with significant systemic racism and bias will likely carry this same bias into implementation" 5. That is the bar any nursing agent is measured against — augmentation with a nurse accountable for the output, and vigilance about whose data trained the system. It aligns cleanly with how the field defines agentic AI in healthcare: autonomy in the mechanics, accountability held by a clinician.

Read operationally, the statement sets two tests a deployment has to pass. The first is accountability: a nurse remains answerable for the note, the assessment, and the handover, which means the agent's output must be reviewable and correctable before it counts — the design the Taiwan team actually built 3. The second is provenance: because biased training data "will likely carry this same bias into implementation" 5, a nurse-facing agent trained on one population's charting can quietly mis-serve another, and a buyer is entitled to ask whose data trained it and on whom it was validated. Neither test is about model quality in the abstract; both are about whether the nurse using the tool can still stand behind what enters the record.

Two claims, two grades of evidence

The single most useful thing a buyer can do is refuse to let these two categories borrow each other's evidence.

ClaimBest current evidenceEvidence grade
Documentation agents save charting timeMulti-hospital before/after, 26-73% time cut on handovers, human-confirmed 3Moderate — real deployment, not randomized
Documentation agents are safe unsupervisedNone; the study that worked disabled risky features and required nurse sign-off 3Not established
Virtual nursing improves documentation completenessObservational, 80% vs 59% admissions, confounded by role design 4Weak — cannot isolate the technology
AI should replace nursing judgmentContradicted by the professional standard 5Rejected

The table's shape is the argument. There is decent evidence that a tightly scoped documentation agent saves time, weak evidence that a virtual-nursing model improves completeness, and no evidence — plus an explicit professional objection — for anything that removes the nurse from the decision. A vendor deck that blurs rows one and three, or that cites the time-savings study to justify autonomy it never tested, is reaching past its data.

How to read this

Three cautions travel with the numbers. First, the strongest result here is a real-world implementation study rather than a randomized trial 3; it shows time saved in practice, which is valuable, but it cannot rule out confounders the way a controlled design would. Second, the completeness figures are observational and confounded by which nurses did which tasks 4, so they measure a staffing model as much as a technology. Third, every one of these studies keeps a nurse reviewing the output, which means none of them tests — or endorses — an unsupervised agent; the burden figures 12 explain the demand, not the safety. The method for grading any of these designs is in our guide on how to read an AI validation study, and the billing-side counterpart to these workflow agents is covered in revenue-cycle and coding agents.

The practical takeaway is a discipline rather than a verdict. When someone offers you an "AI nurse," ask which of the two things it is. If it is a documentation assistant, ask what it is forbidden to touch and who confirms each note — the Taiwan study is strong precisely because it answered both. If it is a virtual-nursing model, ask what the comparison group was doing, because the completeness gain may belong to the role, not the machine. And hold the whole thing against the professional standard: an agent that augments a nurse who stays accountable is on solid ground; anything offered as a replacement is ahead of both the evidence and the profession.

Sources and method

This guide draws on two peer-reviewed studies of nursing documentation burden — a 2025 JMIR Nursing analysis 1 and a 2024 study in the Journal of the American Medical Informatics Association 2 — a 2026 multi-hospital implementation study of an LLM handover tool 3, a 2025 observational study of virtual nursing and task completeness 4, and the American Nurses Association's 2022 position statement on AI in nursing practice 5. Every figure is tied to the primary source cited beside it and was checked live as of July 2026. We revisit this page on a 180-day cycle and whenever a controlled trial of a nurse-facing agent, or an updated professional-body position, is published.

Questions & answers

  • What do AI agents do in nursing workflows?

    Two broad things. Documentation agents draft or restructure nursing notes and handovers from what the nurse provides, saving charting time. Virtual-nursing programs move admission, discharge, and education tasks to a remote nurse, sometimes supported by AI. The documentation use has stronger, more controlled evidence of time savings; the virtual-nursing evidence is largely observational.

  • Do AI documentation tools actually save nurses time?

    The controlled evidence points that way. A 2026 multi-hospital study found an LLM that reformatted nurse-provided input into a structured handover cut per-patient documentation time by roughly a quarter to three-quarters, saving hundreds of nursing hours a month. Crucially, the tool was limited to reformatting and required nurse review of every note.

  • Can an AI agent replace a nurse?

    The professional position is that it augments rather than replaces nursing care. The American Nurses Association's position statement states that AI supports and streamlines expert clinical practice, and every credible deployment keeps a nurse reviewing and confirming the output before it enters the record.

Sources

  1. Evaluating Nurses' Perceptions of Documentation in the Electronic Health Record: Multimethod Analysis. JMIR Nursing. 2025;8:e69651. doi.org/10.2196/69651
  2. Electronic health record system use and documentation burden of acute and critical care nurse clinicians: a mixed-methods study. Journal of the American Medical Informatics Association. 2024;31(11):2637-2650. doi.org/10.1093/jamia/ocae239
  3. Integrating a Large Language Model to Streamline Nursing Handover Documentation Across Multiple Hospitals in Taiwan: Development and Implementation Study. Journal of Medical Internet Research. 2026;28:e81604. doi.org/10.2196/81604
  4. Association of Virtual Nursing and Task Completeness: An Observational Study. SAGE Open Nursing. 2025;11. doi.org/10.1177/23779608251363667
  5. American Nurses Association. The Ethical Use of Artificial Intelligence in Nursing Practice (position statement, adopted 20 December 2022). 2022. www.nursingworld.org/globalassets/practiceandpolicy/nursing-excellence/ana-position-statements/the-ethical-use-of-artificial-intelligence-in-nursing-practice_bod-approved-12_20_22.pdf