Specialties

AI in surgery: a 2026 guide

What artificial intelligence actually does in the operating room today — from polyp detection and anatomy guidance to surgical-phase recognition and the first supervised-autonomy demonstrations — with every capability tied to its primary evidence and its regulatory status. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • AI in the operating room today is mostly perception and prediction — detecting lesions, labeling anatomy, recognizing the phase of an operation, and forecasting instability — not autonomous cutting.
  • The strongest surgical evidence is a procedural one: a 685-patient randomized trial in which real-time computer-aided detection raised colonoscopy adenoma detection from 40.4% to 54.8% (relative risk 1.30) with no added withdrawal time.
  • Autonomy exists as a research milestone: a supervised robot performed laparoscopic small-bowel anastomosis on living tissue, but it remains a demonstration under human oversight, far from a cleared clinical system.
  • A cautionary case runs through the field: the intraoperative Hypotension Prediction Index reported 88% sensitivity 15 minutes ahead, but a re-analysis argues its performance was overstated by selection bias — a simple pressure threshold may explain much of the signal.
  • The cleared reality lags the conversation: in a 2024 device audit, radiology accounted for 74.4% of machine-learning authorizations, leaving operative AI a small, assistive slice — and every serious tool keeps the surgeon responsible for the decision.

Ask what artificial intelligence does in an operating room and the popular image is a robot cutting on its own. The reality in 2026 is quieter and more useful: AI in surgery is mostly a perception-and-prediction layer around a human surgeon — spotting a lesion the eye might miss, labeling which tissue is safe to divide, recognizing which step of an operation is under way, and forecasting a blood-pressure crash before it happens. This guide maps what those systems actually do, ties each capability to its primary evidence, and marks the line between a cleared clinical tool and a research demonstration. As of July 2026.

What "AI in surgery" actually covers

The phrase spans four distinct jobs, and conflating them is the most common error in the field. The foundational review of surgical AI groups the promise and the perils precisely because these jobs carry very different risk 1.

  • Perception — detecting or segmenting structures in a live video feed: a polyp during colonoscopy, the safe and dangerous zones of dissection during a gallbladder removal.
  • Workflow understanding — recognizing which phase of an operation is under way, which enables documentation, scheduling, and real-time prompts.
  • Prediction — forecasting a physiological event, such as intraoperative hypotension, from streaming signals.
  • Autonomy — a robot executing a surgical task with limited human input.

The first three are here today in assistive form. The fourth exists as a milestone rather than a product. Each keeps a surgeon in the loop, which is why the human-in-the-loop design is the defining feature of credible surgical AI rather than an afterthought.

Robotic surgery and surgical AI are different things

A persistent confusion inflates expectations. The widely used surgical robots in operating rooms today are teleoperation systems: a surgeon controls the instruments from a console, and the robot faithfully reproduces their hand movements at smaller scale, with tremor filtered out. That is mechatronics and ergonomics, with little or no artificial intelligence in the surgical decision itself. Surgical AI is the perception, guidance, and prediction layer described above, which may or may not run on a robotic platform. Conflating the two makes autonomous surgery sound closer than it is: a hospital can own advanced surgical robots and use essentially no clinical AI, and much of the evidence in this guide comes from software analyzing ordinary endoscopic or laparoscopic video rather than from any robot 1.

The evidence, capability by capability

The table maps each capability to its strongest primary study and its status. Read every result next to the design that produced it — the discipline our guide on how to read an AI validation study lays out in full. As of July 2026.

CapabilityLandmark evidenceWhat it measuredStatus
Real-time detection (endoscopy)685-patient RCT 2Adenoma detection rateCleared, assistive
Intraoperative anatomy guidanceSegmentation study 3Safe/dangerous dissection zonesResearch → early product
Surgical phase recognitionEndoNet architecture 4Automatic phase labelingResearch → early product
Perioperative predictionWaveform model 6Hypotension 15 min aheadMarketed; contested
Task autonomySupervised robot 5Bowel anastomosis on live tissueResearch demonstration

Detection during procedures: the strongest evidence

The best-evidenced surgical use of AI is a procedural one. In a multicenter randomized trial of 685 patients, a real-time computer-aided detection system that draws a box around suspected lesions on the endoscopy display raised the adenoma detection rate — the share of patients with at least one confirmed adenoma — from 40.4% in the control group to 54.8% with the system, a relative risk of 1.30 (95% CI 1.14–1.45) 2. Adenomas found per colonoscopy rose from 0.71 to 1.07 — an incidence-rate ratio of 1.46 (95% CI 1.15–1.86) — and there was no significant change in withdrawal time, so the gain did not come from simply looking longer 2. This is the archetype for surgical AI that works now: a perception aid, running live, that improves a clinician's yield without taking over the decision. The underlying system was among the first AI tools the FDA authorized for colonoscopy detection. The caveat that keeps this honest is generalization — the trial ran at three expert centers, and whether the same lift appears in community endoscopy, at different baseline detection rates, is the kind of question a local audit answers better than a headline.

Intraoperative anatomy guidance: promising, early

A recurring cause of serious harm in gallbladder surgery is misidentifying anatomy and injuring the bile duct. One response is a computer-vision model trained to label, frame by frame, the zones where dissection is safe and the zones where it is dangerous 3. The work demonstrates that a model can learn a surgeon's notion of a "go" and "no-go" area from annotated video — an early step toward a live safety prompt. It remains developmental: a segmentation result on recorded procedures is a foundation, and the question that matters, whether such prompts change injury rates in live operating rooms, has not yet been answered by a trial.

Surgical phase recognition: the quiet workhorse

Recognizing the steps of an operation from video sounds mundane and is quietly foundational. The EndoNet architecture showed that a convolutional network could learn the phases of a cholecystectomy from visual information alone, seeding a decade of surgical-workflow research 4. Phase recognition powers automatic operative documentation, more accurate operating-room scheduling, and the timing of intraoperative prompts. It rarely makes headlines because it sits underneath other features, but it is among the most mature surgical-AI capabilities and the least controversial, precisely because labeling a step carries little direct patient risk.

Autonomy: a milestone rather than a market

The most striking demonstration is also the most misread. A supervised robot performed laparoscopic small-bowel anastomosis — suturing two segments of intestine — on living tissue 5. Soft-tissue anastomosis is hard because tissue moves and deforms, so this was a genuine advance for autonomy in surgery. It was also a controlled research demonstration under human oversight, on a specific task, far from a cleared system a hospital could adopt. Treat autonomy as a direction of travel rather than an available capability, and be skeptical of any 2026 product that markets full autonomy in the operating room.

Perioperative prediction: a lesson in reading the evidence

Prediction is where surgical AI most needs the reader's discipline. The intraoperative Hypotension Prediction Index was built by relating 3,022 features of the arterial-pressure waveform per cardiac cycle to upcoming hypotension, and reported predicting a low-pressure event 15 minutes ahead with 88% sensitivity and 87% specificity (AUC 0.95) 6. Those are striking numbers. But a later re-analysis argued the performance was overestimated due to selection bias: because of how event and non-event samples were chosen, a mean arterial pressure below 75 mmHg would essentially always predict a hypotensive event, so a simple pressure threshold may explain much of the apparent signal 7. The dispute is unresolved, and that is the point — a headline sensitivity is only as good as the sampling behind it, and sensitivity and specificity figures can flatter a model that has learned little beyond a threshold you already monitor.

What is actually cleared, and what that means

The cleared reality lags the conversation. In an audit of machine-learning devices authorized in 2024, radiology accounted for 74.4% of authorizations, with cardiovascular and neurology next — leaving operative and procedural functions a small slice of the record 10. Our FDA-cleared AI devices tracker holds the fuller breakdown; the pattern for surgery is that most authorized tools are perception aids around a procedure, not autonomous actors within one.

Whether a given surgical tool is regulated as a device turns largely on a single question from the FDA's clinical decision support guidance: can the clinician independently review the basis for the software's recommendation, or must they rely on it? 8 A tool that displays a detection for the surgeon to accept or dismiss sits differently from one that drives an action. As products add real-time guidance and, eventually, autonomy, more of them cross onto the device side of that line — which is why the development principles in Good Machine Learning Practice, issued jointly by the FDA, Health Canada, and the UK MHRA, matter for anyone building in this space 9. For the definition that anchors these distinctions, see software as a medical device and clinical decision support. Device-classification and liability questions are jurisdiction-specific and move quickly; confirm the current status of any tool with your regulatory or compliance counsel before relying on it.

What to ask before adopting a surgical-AI tool

Frame the decision as questions rather than a search for the highest-scoring product. Was the tool evaluated on video or patients that resemble yours — the same procedures, equipment, and case mix — or on a curated benchmark? Is the claim about a proxy, such as a detection or a segmentation accuracy, or about a patient outcome, such as fewer bile-duct injuries or fewer missed lesions? Does the workflow keep the surgeon reviewing and able to override, and how much time does that review add in a real operating room? And is the tool a regulated device with a clearance summary you can read, or a research capability marketed ahead of its evidence? The foundational review's warning still holds: the recurring perils in surgical AI are data quality, hidden bias, and unclear accountability when a tool contributes to a decision, and they are managed by asking these questions before adoption rather than after 1.

How to read these numbers

Four cautions travel with everything above. First, the tasks are not interchangeable: strong randomized evidence for polyp detection says nothing about autonomy, anatomy guidance, or prediction, each of which carries its own risk and its own thin evidence base. Second, a demonstration is a long way from a deployment: the autonomy milestone and the anatomy-guidance work are advances in the laboratory and on recorded video, and the trials that would show they change patient outcomes in live operating rooms have largely not been run. Third, prediction figures can mislead, as the Hypotension Prediction Index dispute shows — always ask what the model was compared against and how its samples were selected 67. Fourth, the surgeon remains accountable: every credible tool here keeps a human making and executing the decision, and the review step is a feature rather than a limitation. For the adjacent imaging field where the evidence and clearances run further ahead, see our AI in radiology guide.

Sources and method

This guide is built on the surgical-AI evidence spine: the foundational review that frames the field 1, the randomized colonoscopy detection trial 2, the intraoperative anatomy-segmentation study 3, the EndoNet phase-recognition architecture 4, the supervised autonomous-anastomosis demonstration 5, and the Hypotension Prediction Index together with its selection-bias re-analysis 67. The regulatory frame draws on the FDA's clinical decision support guidance 8, the Good Machine Learning Practice principles 9, and a peer-reviewed audit of 2024 device authorizations 10. Every figure is tied to the primary source cited beside it. We revisit this page on a 180-day cycle and whenever a new operative-AI authorization or a prospective intraoperative-guidance trial lands. Dates and statuses are current as of July 2026.

Questions & answers

  • Is AI performing surgery on its own yet?

    No. The clearest autonomy milestone is a supervised robot that performed laparoscopic small-bowel anastomosis on living tissue in a research setting, under human oversight. Everyday operative AI is assistive — it detects lesions, labels anatomy, recognizes the phase of an operation, or forecasts instability — and the surgeon remains responsible for every decision and action.

  • What is the best-evidenced use of AI in surgery today?

    Real-time detection during endoscopy has the strongest randomized evidence. In a 685-patient trial, a computer-aided detection system raised the colonoscopy adenoma detection rate from 40.4% to 54.8% (relative risk 1.30) without adding withdrawal time. That is a procedural rather than an operative task, which is part of why its evidence is further along.

  • Are surgical AI tools regulated as medical devices?

    It depends on the function. Whether a tool is a regulated device turns largely on whether the surgeon can independently review the basis for its output, per the FDA's clinical decision support guidance. Detection and guidance systems that drive or could drive a clinical action are more likely to be regulated; confirm the current status of any specific tool with your regulatory or compliance team.

Sources

  1. Hashimoto DA, Rosman G, Rus D, Meireles OR. Artificial Intelligence in Surgery: Promises and Perils. Annals of Surgery. 2018;268(1):70-76. doi.org/10.1097/SLA.0000000000002693
  2. Repici A, Badalamenti M, Maselli R, et al. Efficacy of Real-Time Computer-Aided Detection of Colorectal Neoplasia in a Randomized Trial. Gastroenterology. 2020;159(2):512-520. doi.org/10.1053/j.gastro.2020.04.062
  3. Madani A, Namazi B, Altieri MS, et al. Artificial Intelligence for Intraoperative Guidance: Using Semantic Segmentation to Identify Surgical Anatomy During Laparoscopic Cholecystectomy. Annals of Surgery. 2022;276(2):363-369. doi.org/10.1097/SLA.0000000000004594
  4. Twinanda AP, Shehata S, Mutter D, Marescaux J, de Mathelin M, Padoy N. EndoNet: A Deep Architecture for Recognition Tasks on Laparoscopic Videos. IEEE Transactions on Medical Imaging. 2017;36(1):86-97. doi.org/10.1109/TMI.2016.2593957
  5. Saeidi H, Opfermann JD, Kam M, et al. Autonomous robotic laparoscopic surgery for intestinal anastomosis. Science Robotics. 2022;7(62):eabj2908. doi.org/10.1126/scirobotics.abj2908
  6. Hatib F, Jian Z, Buddi S, et al. Machine-learning Algorithm to Predict Hypotension Based on High-fidelity Arterial Pressure Waveform Analysis. Anesthesiology. 2018;129(4):663-674. doi.org/10.1097/ALN.0000000000002300
  7. Enevoldsen J, Vistisen ST. Performance of the Hypotension Prediction Index May Be Overestimated Due to Selection Bias. Anesthesiology. 2022;137(3):283-289. doi.org/10.1097/ALN.0000000000004320
  8. US Food and Drug Administration. Clinical Decision Support Software — Guidance for Industry and FDA Staff. (accessed July 2026). www.fda.gov/regulatory-information/search-fda-guidance-documents/clinical-decision-support-software
  9. US FDA, Health Canada, and UK MHRA. Good Machine Learning Practice for Medical Device Development: Guiding Principles. October 2021. www.fda.gov/medical-devices/software-medical-device-samd/good-machine-learning-practice-medical-device-development-guiding-principles
  10. Almarie B, et al. Machine Learning-Enabled Medical Devices Authorized by the US Food and Drug Administration: Regulatory Characteristics, Predicate Lineage, and Transparency Reporting. Biomedicines. 2025;13(12):3005. doi.org/10.3390/biomedicines13123005