Specialties

AI in pharmacy: a 2026 evidence guide

What the published record actually supports for AI in pharmacy — where dispensing automation has measured safety gains, why interruptive drug-interaction alerts are overridden roughly nine times in ten, and what a language model can and cannot be trusted to do with a drug question. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • The strongest pharmacy evidence is mechanical rather than cognitive: robotic dispensing systems have cut dispensing-error rates several-fold and dropped per-prescription handling time from about a minute to roughly twenty seconds in controlled hospital studies.
  • Interruptive drug-drug interaction alerts are overridden about 90% of the time (pooled across 16 studies), so the real frontier is fewer, smarter, patient-specific alerts rather than more of them.
  • Language models answer pharmacy-exam questions well — GPT-4 scored 83.5% and 87% on two NAPLEX practice sets — but accuracy drops from recall to application, reproducibility is unreliable, and drug-specific hallucination remains a safety risk.
  • The FDA runs a dedicated AI-for-drug-development program and describes AI across pharmacovigilance; the ASHP's 2025 statement puts pharmacists in charge of designing, validating, and overseeing any AI in the medication-use process.
  • As of July 2026, no published evidence supports an unsupervised AI making a dispensing or clinical-pharmacy decision; every credible deployment keeps a pharmacist accountable for the output.

Pharmacy is where several kinds of AI meet in one workflow, and where the gap between marketing and evidence is easiest to see. A hospital pharmacy runs a robot that picks and packs, a decision-support layer that screens every order for interactions, and — increasingly — a language model that a technician or pharmacist asks for a quick answer. Each of these has a different maturity, a different failure mode, and a different weight of published evidence behind it. This guide walks the three, holds each to what the record supports, and marks where a pharmacist still has to be the last set of eyes. As of July 2026.

Where the evidence is strongest: dispensing automation

The most convincing pharmacy AI is the least glamorous. Robotic dispensing and automated storage-and-retrieval systems have been measured in real hospital pharmacies, and the numbers move in the right direction. A controlled study of a robotic dispensing system reported that prevented dispensing errors fell from 0.204% to 0.044% and unprevented dispensing errors from 0.015% to 0.002%, while the median dispensing time per prescription dropped from 60 to 23 seconds 3. Those reductions held after the system had settled into routine use alongside pharmacy staff, which matters — many technology gains erode once the novelty passes.

The broader literature agrees. A 2026 systematic review and meta-analysis of hospital pharmacy automation across G20 countries found automation associated with fewer dispensing-related errors, reporting a risk ratio of 3.52 favouring automated over manual dispensing, with wrong-drug and wrong-dose errors the most reduced, alongside efficiency and staff-satisfaction gains 5. A 2025 evaluation of an automated drug-retrieval cabinet paired with a robotic dispensing system in a large central pharmacy documented the same pattern of reduced filling errors and recovered pharmacist and technician time 4.

Read carefully, this is a story about a well-scoped task. Dispensing is physical, repeatable, and verifiable: the right drug, the right strength, the right count, checked against a barcode. Automation excels precisely because the task has a single correct answer that a machine can hit more consistently than a tired human at hour ten of a shift. That is the template worth remembering as the rest of this guide moves toward tasks where the "right answer" is fuzzier and the machine's advantage narrows.

One boundary is worth marking before moving on. Dispensing automation addresses the dispensing step — pulling, counting, labelling, checking — while a large share of preventable medication harm originates earlier, at prescribing, or later, at administration. A flawless robot fills the wrong drug perfectly if the order was wrong. The efficiency and safety gains in the G20 review are real, but they are gains at one link in a longer chain 5; the value of pharmacy AI further upstream, in catching a prescribing error before it is ever dispensed, is exactly what the interaction-alert problem below is about.

The alert-fatigue problem AI inherits

Every pharmacy already runs a form of AI-adjacent automation: the clinical decision support layer that screens orders for drug-drug interactions, allergies, and dosing problems. The trouble is well quantified. A 2024 systematic review and meta-analysis of 16 studies found the pooled drug-drug interaction alert override rate was 90% (95% CI 85-95%), with individual studies ranging from 55% to 98% 1. In other words, roughly nine interaction alerts in ten are dismissed without changing the order.

Some of those overrides are appropriate — the alert was a false positive, or the interaction was clinically irrelevant for that patient. The deeper problem is what a wall of low-value alerts does to the high-value ones. When most warnings are noise, the signal that should stop a dangerous order gets clicked through with the rest. The review traces the cause to alerts that fire on too broad a screening interval and ignore patient-specific characteristics such as renal function or dose 1.

This reframes where AI can genuinely help. The win is a smaller, smarter set of alerts rather than a new layer of them — models that suppress the irrelevant and surface the patient-specific, so a pharmacist sees ten warnings that matter for every hundred that today mostly do nothing. Any vendor promising "more comprehensive" interaction checking is describing the problem, not the solution. The measure of a good tool here is how much it reduces interruptions while holding onto the ones that change care.

Language models and drug knowledge

The newest entrant is the large language model, and its pharmacy report card is genuinely mixed. On structured knowledge it performs well: one evaluation found GPT-4 answered 87% of a McGraw Hill NAPLEX practice set and 83.5% of an RxPrep set correctly, far above GPT-3.5 at 68% and 60% 2. It was strongest on adverse-drug-reaction questions, at 96% 2.

Three caveats sit under those headlines, and each one shapes safe use. First, question type matters: every model tested did markedly worse on select-all-that-apply items than on single-answer ones — GPT-4 fell from 87% to 73% 2 — which is the format closest to real clinical reasoning, where several things are true at once. Second, application lags recall. A 2024 evaluation of language models on critical-care pharmacy assessments found GPT-4 most accurate overall at 71.6%, but every model performed better on knowledge recall than on knowledge application, and reproducibility across repeated runs was unreliable 8. A tool that gives a different answer to the same question on Tuesday than on Monday cannot anchor a medication decision.

Third, and most important for a drug context, is fabrication. A clinical language model that hallucinates a dose, an interaction, or a contraindication is more dangerous in pharmacy than almost anywhere else, because the output looks like exactly the kind of precise, confident statement a pharmacist is trained to trust. The safe posture follows directly: a model answer is a draft to be checked against a primary reference, never the reference itself.

Where does that leave a pharmacist who wants to use these tools? On safe ground for language tasks that a human then verifies: drafting a patient-friendly explanation of a regimen, translating instructions into another language, condensing a long medication history into a starting point, or surfacing candidate interactions to confirm against a primary source. The unifying rule is that the model drafts or reorganizes and the pharmacist decides — the same division of labor that made constrained documentation tools safe elsewhere in healthcare. What a model must never be is the final authority on a dose, an interaction, or a contraindication. For how these benchmark numbers translate — and fail to translate — into clinical reliability, see our LLM medical benchmark results tracker and the method in how to read an AI validation study.

Regulators are watching the drug lifecycle beyond the device

AI in pharmacy also reaches upstream, into how drugs are developed and how their safety is monitored after approval. The FDA runs a dedicated Artificial Intelligence for Drug Development program and has issued a discussion paper (2023, revised 2025) together with draft guidance describing how AI and machine learning are used across the drug lifecycle — including pharmacovigilance, where the agency describes AI supporting case processing, case evaluation, and case submission of individual adverse-event reports 6. Confirm the live status of any FDA guidance before relying on it; drafts change.

The practical signal for a pharmacy leader is that adverse-event surveillance — reading free-text reports, spotting a possible causal signal, routing a case — is an area regulators expect AI to touch, under human review. It is a natural fit: high volume, pattern-heavy, and consequential enough that a person still signs off.

The professional standard

Pharmacy, like nursing, has put its position in writing. The ASHP Statement on Artificial Intelligence in Pharmacy (2025), which supersedes its 2020 statement, holds that pharmacists should lead the design, implementation, validation, and maintenance of AI tools affecting the medication-use process, and calls for human oversight, auditability, and transparency in those tools 7.

Read operationally, that sets two tests any pharmacy AI has to pass. The first is accountability: a pharmacist remains answerable for the dispensed product and the clinical recommendation, which means the tool's output must be reviewable and correctable before it counts — the human in the loop as a design feature. The second is stewardship: pharmacists, not vendors, should own the selection, validation, and monitoring of the tool, because they are the ones who will answer for a medication error. Neither test is about model sophistication in the abstract; both are about whether the pharmacist using the tool can still stand behind what reaches the patient.

Three capabilities, three grades of evidence

The single most useful discipline is to refuse to let these categories borrow each other's credibility.

CapabilityBest current evidenceEvidence grade
Dispensing automation reduces errors and timeControlled hospital studies; errors down several-fold, handling time 60→23s 3; meta-analysis risk ratio 3.52 5Moderate-to-strong
Smarter interaction alerts cut fatigue~90% of current DDI alerts overridden; patient-specific tuning is the stated fix 1Problem well-established; solutions early
Language models answer drug questionsGPT-4 83.5-87% on NAPLEX practice sets 2; 71.6% on critical-care items, recall > application, low reproducibility 8Weak for autonomous use
AI in pharmacovigilanceRegulator-described uses under human review 6Emerging, guidance-stage
Unsupervised clinical-pharmacy decisionsNone 7Not established

The shape of the table is the argument. There is decent evidence that a well-scoped machine dispenses more safely than a person, a well-established problem that current interaction alerts fire too often, and no evidence — plus an explicit professional objection — for anything that removes the pharmacist from the decision.

How to read this

Three cautions travel with the numbers. First, the dispensing results are the strongest here because the task is verifiable; do not let a robot's proven reliability at counting tablets transfer, in a vendor's pitch, to a language model's unproven reliability at reasoning about a drug. Second, benchmark accuracy marks a ceiling rather than a floor: an 87% on practice questions was measured on clean, single-turn items, and real practice is messier, multi-part, and higher-stakes. Third, the professional standard is unambiguous that a pharmacist owns the output, which means none of these tools is validated — or endorsed — for unsupervised use.

Because dispensing, compounding, and clinical-pharmacy recommendations carry direct liability and are bound by scope-of-practice and licensing rules that differ by jurisdiction, confirm any AI-enabled workflow with your pharmacy compliance and legal counsel before deployment — clearance or CE status is not the same as a settled liability position. Our guide on liability when clinical AI errs covers the terrain, and the billing-side counterpart to medication and order workflows is in revenue-cycle and coding agents.

Sources and method

This guide draws on a 2024 systematic review and meta-analysis of drug-drug interaction alert override rates 1, a 2024 evaluation of large language models on pharmacist-licensure practice questions 2, a controlled study of a robotic dispensing system 3, a 2025 evaluation of an automated drug-retrieval and robotic dispensing system 4, a 2026 systematic review and meta-analysis of hospital pharmacy automation 5, the FDA's Artificial Intelligence for Drug Development program materials 6, the ASHP's 2025 professional statement 7, and a 2024 evaluation of language-model performance on critical-care pharmacy assessments 8. Every figure is tied to the primary source cited beside it and was checked live as of July 2026. We revisit this page on a 180-day cycle and whenever a controlled trial of a pharmacy AI tool, a revised FDA guidance, or an updated professional statement is published.

Questions & answers

  • What does AI actually do in pharmacy today?

    Three things, at very different levels of maturity. Dispensing automation — robotic cabinets and dispensing systems — has the strongest evidence, with controlled hospital studies showing several-fold reductions in dispensing errors and large time savings. Clinical decision support, such as drug-drug interaction checking, is widespread but undermined by alert fatigue. Language models can answer drug questions and draft documentation, but their accuracy is uneven and they require pharmacist review before anything reaches a patient.

  • Are drug-drug interaction alerts useful if clinicians override most of them?

    The alerts catch real problems, but a pooled override rate near 90% means most fire without changing a decision, and the volume risks burying the alerts that matter. The improvement path is not more alerts but patient-specific ones — tuned to renal function, dose, and context — so a smaller number of higher-value warnings reach the pharmacist.

  • Can a language model be trusted to answer a drug-information question?

    Not on its own. GPT-4 scores in the mid-80s on NAPLEX practice questions, but performance falls on application versus recall, answers can change between runs, and drug-specific fabrication is a documented risk. Treat a model answer as a draft a pharmacist verifies against a primary reference, never as the reference itself.

Sources

  1. Felisberto M, dos Santos Lima G, Celuppi IC, et al. Override rate of drug-drug interaction alerts in clinical decision support systems: A brief systematic review and meta-analysis. Health Informatics Journal. 2024;30(2). doi.org/10.1177/14604582241263242
  2. Angel MC, Rinehart JB, Canneson MP, Baldi P. Large Language Models and the North American Pharmacist Licensure Examination (NAPLEX) Practice Questions. American Journal of Pharmaceutical Education. 2024;88(11):101294. doi.org/10.1016/j.ajpe.2024.101294
  3. Kwon M, et al. Evaluating the safety and efficiency of robotic dispensing systems. Journal of Pharmaceutical Health Care and Sciences. 2022;8:23. doi.org/10.1186/s40780-022-00255-w
  4. Evaluating the impact of an automated drug retrieval cabinet and robotic dispensing system in a large hospital central pharmacy. American Journal of Health-System Pharmacy. 2025;82(1):32-40. doi.org/10.1093/ajhp/zxae225
  5. Alshmemri MA, Mady FM, Sadek EM, Hussein AK, et al. Impact of pharmacy automation on dispensing efficiency, medication safety, and user satisfaction in hospital pharmacies: A systematic review and meta-analysis of G20 countries. Pharmacia. 2026;73:e190314. doi.org/10.3897/pharmacia.73.e190314
  6. US Food and Drug Administration. Artificial Intelligence for Drug Development (CDER). Accessed July 2026. www.fda.gov/about-fda/center-drug-evaluation-and-research-cder/artificial-intelligence-drug-development
  7. ASHP Statement on Artificial Intelligence in Pharmacy. American Journal of Health-System Pharmacy. 2025;82(19):e853-e862. doi.org/10.1093/ajhp/zxaf107
  8. Evaluating accuracy and reproducibility of large language model performance on critical care assessments in pharmacy education. Frontiers in Artificial Intelligence. 2024;7:1514896. doi.org/10.3389/frai.2024.1514896