The kit
The clinical AI evaluation kit
Four things to take into the room where a clinical AI tool is bought: the questions to ask the vendor, a checklist that reconciles the published evaluation frameworks, the red flags in a reported performance figure, and a governance charter with an intake form. Nothing here is gated and nothing here is for sale.
Open to read · nothing to fill in · every claim sourced
Contents
- 00How to use this
- 01The vendor question list
- 02The evaluation checklist
- 03Red flags in a reported figure
- 04A governance starter charter
- 05What this kit cannot tell you
- 06Sources
Last checked 13 August 2026. Figures carry the date they were true.
How to use this
Four artifacts, in the order you will need them
Buying a clinical AI tool involves three separate judgements that are usually collapsed into one: whether the evidence is any good, whether it is evidence about your patients, and who is accountable once the tool is live. The kit splits them.
On why this is a live problem rather than a coming one: in the American Medical Association’s survey of 1,692 US physicians, fielded 15 January to 2 February 2026, 81% reported some awareness or use of AI (n = 1,342 for that item). The AMA notes that the 2026 wave counts qualified partial responses where earlier waves counted only complete ones, so it is not a like-for-like comparison with previous years [39].
This is orientation and general guidance, not legal, clinical or regulatory advice. It is not accredited, certified or endorsed by any body, and it does not substitute for your own compliance, legal and clinical review. Where the published sources disagree, the disagreement is shown rather than resolved. Where something is genuinely unknown, it is in what this kit cannot tell you.
Artifact 01
The vendor question list
Sixteen questions, in the order the process runs. The first four are answered by documents and should be asked before anyone books a demo; the middle six are about the number on the slide; the last six decide what year two looks like.
16 questions3 stagesGood answer and deflection on each
Before the demo, from the paperwork
These four are answered by documents, not by a person. Ask for the documents and read them before the meeting, because every later question depends on what they say.
Before the demo · 1
“Read me the indications-for-use statement, word for word.”
The cleared indication is the only sentence that binds. It names the patients, the modality, the setting and the role the output plays, and it is routinely narrower than the sales conversation. Everything after this question is measured against it. [14][16]
A good answer
They read the statement from the clearance letter or the labelling, then say plainly which parts of your intended use fall outside it and what that means for you.
A deflection
A paraphrase, a slide bullet, or a jump to the clinical problem you have. If nobody in the room can produce the sentence, the meeting is about a product that has not been described yet.
Before the demo · 2
“Which regulatory pathway, and what did it establish?”
The pathways prove different things. A 510(k) establishes substantial equivalence to an already-marketed device; it is not a finding that the device improves outcomes at your hospital. De Novo and PMA carry different evidence. Some tools are not devices at all. [12][14]
A good answer
The pathway, the submission number, the date, and an unforced statement of what the clearance did and did not test.
A deflection
“FDA-approved” used as a single undifferentiated word, or “FDA-cleared” for a product that was never reviewed because it sits inside a non-device exemption.
Before the demo · 3
“Is there a Predetermined Change Control Plan, and what does it let you change without telling us?”
A PCCP is the mechanism by which an authorised device can be modified after clearance without a new submission. It is legitimate and it is public. What matters to you is its scope: what the vendor may retrain, on what data, and what they owe you when they do. [13][37]
A good answer
They hand you the plan, walk through the modification protocol and the impact assessment, and tell you how a change reaches your site and how you would know.
A deflection
“The model improves continuously.” That sentence describes either a PCCP the vendor will not show you or a change process with no regulatory basis. Both are answers you need.
Before the demo · 4
“If this runs inside our certified EHR, produce the source attributes.”
In the United States, certified health IT that supplies predictive decision support must make a defined set of source attributes available to users. They cover intended use, training data, performance, validation and maintenance. You do not have to negotiate for them. [16]
A good answer
The complete attribute set, supplied as a document, without a request for a non-disclosure agreement first.
A deflection
“That is proprietary.” Some of it may be. The certified attribute set is not, and a vendor who does not know the requirement exists has told you something about their regulatory function.
In the demo
Six questions about the number on the slide. Ask them in this order: each one narrows what the previous answer could have meant.
In the demo · 1
“Where was this validated outside the site that built it? How many sites, whose patients, what years?”
Performance measured on held-out data from the training source tells you how well the model learned that source. External validation on different sites, scanners and eras is the only evidence that travels, and it is rare enough that its absence is the default. [1][31]
A good answer
Named sites, patient counts, the calendar years of the data, the case mix, and the performance at each site separately rather than pooled.
A deflection
A single pooled figure across “multiple health systems”, or cross-validation described as if it were external validation.
In the demo · 2
“What is the performance at the threshold we will actually run?”
AUROC summarises every possible threshold, including the ones nobody would use. Deployment happens at one operating point. Sensitivity, specificity, alert rate and workload all move with it, and the vendor chose the one on the slide. [4][37]
A good answer
A table of sensitivity and specificity at several thresholds, the threshold they recommend, the reason, and the alert volume that follows from it.
A deflection
Only an AUROC, or a threshold that is “configurable” with no default and no guidance on how to configure it.
In the demo · 3
“At our prevalence, how many of the flags will be right?”
Positive predictive value depends on how common the condition is in the patients you actually see. A tool validated on an enriched sample can post excellent sensitivity and specificity and still produce mostly false alarms at a realistic base rate. This is arithmetic, and it can be done in the meeting. [32]
A good answer
They ask for your prevalence and do the calculation with you, or they already have the curve of PPV against prevalence.
A deflection
Accuracy quoted as a single percentage, or a refusal to discuss prevalence because “it varies by site”. It does vary by site; that is the point of the question.
In the demo · 4
“Show me the calibration curve.”
Discrimination tells you whether the model ranks sicker patients above healthier ones. Calibration tells you whether a predicted 20% risk corresponds to a real 20% event rate. A model can rank well for years while its probabilities drift out of true, and calibration is the metric most often left out. [33][38]
A good answer
A calibration plot with the observed-against-predicted line, the calibration slope and intercept, and the date and population it was measured on.
A deflection
Confusion between calibration and accuracy, or a promise that the model is “calibrated to your data” during implementation with no description of how that is checked afterwards.
In the demo · 5
“Break the performance down by the subgroups in our population.”
Overall performance can hold while a subgroup fails. The relevant subgroups are the ones in your catchment: age bands, sex, the languages your patients speak, skin tone where imaging is involved, comorbidity burden, the sites and scanners you run. [1][36][37]
A good answer
Subgroup results with the sample size in each cell, and an honest statement of which subgroups were too small to say anything about.
A deflection
A fairness statement with no numbers, a single aggregate “bias audit passed”, or subgroup results with the denominators removed.
In the demo · 6
“How does it fail, and does it know when it is out of its depth?”
A model given an input unlike anything in its training data still returns an answer. What you need to know is whether the system detects that condition and abstains, degrades, or flags it — and what the clinician sees when it does. [1][16]
A good answer
A described out-of-distribution behaviour, a rejection or low-confidence path, and screenshots of what the user sees. Ideally, examples of inputs that trip it.
A deflection
“It handles edge cases.” Ask for three examples of inputs on which the vendor would not want the output trusted. A vendor who cannot name any has not looked.
Before you sign
Six questions the clinical team will not ask and the procurement team will not know to ask. They decide what happens in year two.
Before you sign · 1
“What is monitored after go-live, by whom, at what cadence, and what do we see?”
Model performance decays as the population, the coding practice and the upstream systems move. Approval is the start of the obligation, not the end of it. Somebody has to be watching discrimination, calibration and subgroup performance on your data, on a schedule. [15][16]
A good answer
A named monitoring plan with metrics, cadence, thresholds, an owner on each side, and a dashboard you can see without asking.
A deflection
Monitoring described as an optional professional-services engagement, or as something the vendor does internally and reports on annually.
Before you sign · 2
“What triggers a retrain, how are we told, and can we decline it?”
A model that changes underneath a clinical workflow is a new model. You need to know what causes a version change, what notice you get, whether you can stay on a version, and how you would revalidate. [13]
A good answer
A versioning policy, a notification period, a release note that states what changed and what it did to performance, and a documented way to pin or roll back.
A deflection
“Updates are seamless.” Seamless means you will not be told.
Before you sign · 3
“What is the alert burden, in alerts per hundred patients per day?”
The cost of a decision-support tool is mostly paid in attention. Alert volume and override rate decide whether clinicians read it in month six. Reference sites can supply both numbers, and vendors rarely volunteer them. [5]
A good answer
Real deployment numbers from named sites: alerts per unit of clinical activity, override rate, and how both moved over the first year.
A deflection
Satisfaction scores instead of alert rates, or pilot-period numbers presented as steady state.
Before you sign · 4
“Name three customers who stopped using it, and tell me why.”
Reference customers are selected. Churned customers are the evidence you are not being shown, and the reasons are usually about integration, alert burden, or a performance gap that only appeared locally.
A good answer
Names, or at least the reasons, given without visible discomfort. A vendor with a mature product knows exactly why they lose sites.
A deflection
“We have never had a customer leave.” For any product old enough to buy, that is either untrue or means nobody has been using it.
Before you sign · 5
“What happens to our data, and is any of it used to train your models?”
Answers vary from “nothing leaves the boundary” to “de-identified data trains the shared model”, and the contract, not the demo, decides which. This also determines what your privacy office has to sign and what you have to tell patients.
A good answer
A clear statement of what leaves your environment, where it is processed, how long it is retained, what secondary use is permitted, and where in the contract that is written down.
A deflection
“Everything is HIPAA compliant.” That is a claim about a framework, not a description of a data flow.
Before you sign · 6
“If we turn it off, what do we keep?”
Exit terms decide how reversible this decision is. Model outputs written into the record, the audit trail, and the monitoring history all have to survive the end of the contract, because they may be needed years later.
A good answer
Named export formats, a stated retention period after termination, and confirmation that outputs already written to the record remain readable without the vendor.
A deflection
An exit clause nobody in the room has read.
Artifact 02
The evaluation checklist
One page, reconciled from frameworks that were written for different readers. Reporting guidelines tell an author what to disclose; appraisal tools score a paper that is already published; consensus guidelines describe a lifecycle. None of them was written for the person deciding whether to buy.
35 items6 stagesReconciled from 9 frameworks
What was reconciled
Item counts are taken from each framework’s own publication, checked 13 August 2026. They are not comparable with one another: an “item” means a reporting element in one and a scored question in another, which is the first reason a reader cannot simply pick the longest one.
| Framework | Size | Written for | What it does not cover |
|---|---|---|---|
| FUTURE-AI [1] | 6 guiding principles, 30 recommendations (2025). Built by a consortium of 117 experts across 50 countries. | Developers, over the whole lifecycle from design to post-deployment monitoring. | Procurement. The words purchase, procurement and tender do not appear in it, and most of its 30 recommendations describe what a developer should do rather than anything a buyer can check from outside. |
| PROBAST+AI [2] | 34 signalling questions in two parts — 16 for model development, 18 for model evaluation — across 4 domains (2025). | Anyone appraising a model. Uniquely, its own stakeholder table names healthcare professionals verifying a model “before purchasing or using” it. | It assesses the study, not the deployment. It will tell you whether a result is at risk of bias; it will not tell you what the tool does to your clinic. |
| APPRAISE-AI [3] | 24 items in 6 domains, scored out of 100 points; validated against 28 machine-learning sepsis-prediction studies (2023). | Investigators, reviewers, editors and funders scoring a published study. | Its authors exclude implementation by design: it “is not intended to evaluate feasibility and other ethical considerations that are essential to clinical implementation, such as ease of use, interoperability, and privacy concerns.” |
| TRIPOD+AI [4] | 27 items (2024). Supersedes TRIPOD 2015, which its authors say should no longer be used. | Authors reporting the development or evaluation of a clinical prediction model. | It governs disclosure, not fitness. A study can satisfy every item and still describe a model that is wrong for your population. |
| DECIDE-AI [5] | 27 items — 17 AI-specific ones containing 28 subitems, plus 10 generic (2022). | Investigators reporting early-stage, small-scale live clinical evaluation. | It is deliberately scoped to one stage: after offline validation, before large trials. Most tools on sale have no study of this kind at all. |
| CLAIM [6][11] | 44 items in the 2024 update; 42 in the 2020 original. Neither version states its own total in a sentence. | Authors and reviewers of medical-imaging AI studies. | Imaging only, and reporting only. A separate review scored the same studies out of 53 by counting subitems, which is one reason circulating item counts disagree. |
| CONSORT-AI [7] | 14 new AI-specific items, added on top of the core CONSORT 2010 checklist (2020). It is not a 14-item checklist. | Authors reporting a randomised trial of an AI intervention. | It applies only where a randomised trial exists. For most tools being sold, one does not. |
| MI-CLAIM [8] | 6 parts. The paper never states an item count; counting the authors’ own checklist file gives 18 rows plus a 4-tier transparency selection. | Algorithm designers, repository managers, manuscript writers, editors and model users. | Retrospective modelling studies. Any total you see quoted for it is arithmetic somebody else did. |
| MINIMAR [9] | 4 components, 21 named reporting fields in the authors’ own table. No total is stated. | Research scientists and medical practitioners, and journals. | Published explicitly as “a starting point for a broader community discussion”, with four authors and no consensus process. Its strength is demographic reporting, which most of the others under-specify. |
The reconciliation rule is simple and worth stating, because it is a judgement rather than a finding: an item is in this checklist if at least one of the frameworks requires it and a buyer could actually check it before signing. Items that only an author or a peer reviewer can verify were dropped. Items that only appear in one framework are kept and marked, so you can see how thin the agreement is in places.
A · Is this evidence about your patients?5 items
B · How was it tested?6 items
C · What was actually reported?6 items
D · What does it do to the workflow?5 items
E · What happens after go-live?6 items
F · Regulatory and contractual7 items
Ticks are yours alone. Nothing on this page is recorded, stored or sent anywhere, and reloading clears them.
Artifact 03
Red flags in a reported figure
These are about the number itself, not the study behind it. The study-design version — ten flags, each anchored to a documented published failure — is a separate guide, and this table does not repeat it.
9 flagsOne worked calculationThe study-design version
Each row is a property of the figure you are being shown. One flag is a question; three together are a pattern.
| What you are shown | Why it will not hold | Ask for |
|---|---|---|
| A single accuracy percentage | Accuracy tracks the commonest answer. Where a condition affects 1 patient in 100, a model that says “no” every time is 99% accurate and clinically worthless. | Sensitivity, specificity, and positive predictive value at your own prevalence. |
| An area under the curve, and nothing else | It averages performance across every threshold, including ones nobody would deploy. You will run exactly one, and the vendor chose which one to show. [4] | A table of sensitivity and specificity across thresholds, plus the recommended operating point and why. |
| Discrimination with no calibration | Ranking patients correctly is not the same as predicting the right probability. Miscalibrated outputs are described in the methods literature as potentially misleading and harmful for clinical decision-making. [33] | A calibration curve with slope and intercept, and the population and date it was measured on. |
| A figure with no confidence interval | A point estimate hides its own precision. The same 0.92 from a few dozen outcome events and from several thousand are different claims about the world. [4] | The interval, and the number of outcome events behind it. |
| One number pooled across sites | Pooling averages away the site where it failed. Differences between sites are the thing you are trying to measure, because you are about to become a new one. [1] | Per-site results with the sample size at each. |
| An internal-validation figure presented as performance | Measured on held-out data from the same source, it describes how well the model learned that source. Across 516 studies of AI for diagnostic analysis of medical images, only 6% (31 studies) performed external validation. [31] | Evaluation on data from a site, scanner or era that took no part in development. |
| Overall performance with no subgroup breakdown | Aggregate performance can hold while a subgroup fails. A care-management algorithm applied to millions of patients was found to rank Black patients as lower risk than equally sick White patients, because it optimised on cost as a proxy for need. [36] | Subgroup results with denominators, for the subgroups in your population. |
| A claim of parity with clinicians from a retrospective single-site study | The design cannot carry the sentence. Of 81 studies comparing deep learning against clinicians in medical imaging, 61 concluded the algorithm was comparable or superior, while most non-randomised studies were not prospective and were at high risk of bias. [34] | The comparator, the design, and a claim proportionate to both. |
| A number with no date and no model version | The model that produced the figure may not be the model on sale. Where an authorised Predetermined Change Control Plan exists, specified modifications can be implemented without a new marketing submission. [13] | The version the figure describes, the date it was measured, and whether it was re-measured after the most recent change. |
The one calculation worth doing in the room
For the full ten, each tied to a documented published failure, see ten red flags of an overfit model claim, and for the study-reading method, how to read an AI validation study.
Artifact 04
A governance committee starter charter
Until May 2026 no authority published anything at this level of detail. That changed, and this section says so before it says anything else: there are now templates, they are worth reading, and this charter is offered as one filled-in position rather than as a substitute for them.
Three things shipped in the second quarter of 2026. The Coalition for Health AI released eight governance playbooks on 27 May 2026, whose second domain covers organisational structures with four baseline controls — critical roles, a formal committee, an intake and monitoring process, and escalation protocols — plus a worked three-tier escalation example and case studies from named health systems [21]. The Health Sector Coordinating Council’s implementation guide, published in May 2026, carries an actual committee charter template, a RACI matrix, a use-case risk-scoring form and a board reporting template as appendices, though it is scoped to the cybersecurity dimensions of governance [23]. And the Joint Commission announced a Responsible Use of AI in Healthcare certification on 1 June 2026 [22].
What none of them supplies is a filled-in answer. Every one defers membership, quorum and decision rights to “as defined by your organisation”; the charter template with a quorum field leaves the number blank; and the published descriptions of committees that actually run are case reports of one institution, not templates [28][29]. That is the gap this fills: a specific, arguable default for the five decisions everyone else leaves open. It is also, still, entirely voluntary — see what this kit cannot tell you.
The layers above and below this one are already covered elsewhere and are not repeated here. For stated values, the National Academy of Medicine’s code of conduct of 14 July 2025 sets out ten principles and six commitments [24], and the World Health Organization’s 2021 guidance sets out six [27]. For the sequence of decisions a tool passes through, the Health AI Partnership publishes eight key decision points across four phases with 31 best-practice guides [25], and the American Medical Association’s governance toolkit of 8 May 2025 sets out eight steps, the fifth of which is defining intake and vendor evaluation [26].
10 seats3 tiers6 escalation triggers14 intake fields
Purpose, in one paragraph
The [committee name] reviews, approves, restricts and retires every artificial intelligence tool used at [organisation] that touches patient care, clinical documentation or clinical operations. It maintains the inventory of those tools, sets the evidence required at each risk tier, holds the local-validation gate, owns post-deployment monitoring, and reports to [the board committee] at least [quarterly]. It has the authority to stop a tool. A body that can review but cannot restrict or retire is not this committee.
Membership, and what each seat can stop
Ten seats. The third column is the part usually left out: a seat with no defined authority attends rather than governs.
| Seat | What it brings | What it can stop on its own |
|---|---|---|
| Executive sponsor (chair) | Budget authority and a direct line to the board. Owns the standing report upward. | Any deployment. The chair is the only seat that can approve a tier 1 tool going live. |
| Clinical lead for the affected service | Whether the workflow described in the submission is the workflow that exists. | Go-live in their own service. A tool cannot be imposed on a service over its clinical lead. |
| Quality and patient safety | The route from an AI-related incident into the existing safety reporting system. | Continued operation after an unresolved safety signal. |
| Clinical informatics | Integration reality, data pipelines, and the local validation itself. | Deployment without a completed local performance check. |
| Privacy and information security | Data flows, retention, secondary use, and the vendor’s security posture. | Any data leaving the boundary on terms the office has not approved. |
| Legal, compliance or regulatory | Regulatory fit, contract terms, disclosure obligations, and liability exposure. | Signature. Nothing is contracted over a compliance objection. |
| Health equity or population health | Which subgroups matter here, and whether the evidence covers them. | Deployment where a subgroup in the local population has no performance evidence. |
| Nursing and allied health | The cost of the tool in attention, for the staff who absorb most of it. | None. Advisory, but the chair records a dissent in the minute. |
| Patient or community representative | What the organisation is prepared to say to the people the tool is used on. | None. Advisory, with the same recorded-dissent right. |
| Data science or model owner | Reads the vendor’s evidence critically and runs the local check. | None. Presents; does not vote on their own submission. |
Tiers, evidence, and who decides
Uniform scrutiny wastes the committee on schedulers and under-examines diagnostics. Tier by proximity to a clinical decision, and let the tier set both the evidence bar and the monitoring cadence.
| Tier | What is in it | Evidence required | Who decides | Monitoring |
|---|---|---|---|---|
| Tier 1 — drives a clinical decision | Diagnosis, triage, risk scores that change management, autonomous or near-autonomous output. | External validation, local performance check on your own data before go-live, subgroup results, calibration, a monitoring plan. | Full committee. Chair approves. Quorum: chair, clinical lead, informatics, safety, compliance. | Monthly for the first quarter, then quarterly. Board hears about it by name. |
| Tier 2 — informs a clinical decision | Documentation, summarisation, drafting, prioritisation of a queue a human still works through. | Vendor evidence read in full, a local pilot with an error-rate audit, a named human review step. | Full committee. Simple majority. Quorum: five seats including clinical and compliance. | Quarterly, plus an incident route. |
| Tier 3 — no clinical decision | Scheduling, coding support, back-office, staffing, supply. | Privacy and security review, a stated owner, entry in the inventory. | Delegated to a two-seat subgroup. Reported to the committee, not debated by it. | Annual review. Escalates on incident. |
What forces an item back onto the agenda
Escalation is what separates a governance body from a reading group. Each trigger is written so that it fires without anyone deciding it should.
| Trigger | What happens |
|---|---|
| A monitored metric crosses its pre-agreed threshold | Automatic item at the next meeting. The threshold is set at approval, in writing, by the committee — not by the vendor. |
| A subgroup’s performance falls below the deployment-wide floor | Restrict to the populations where evidence holds, within the same working week, pending review. |
| An AI-related safety event is reported | Into the existing safety system on the existing timescale. The committee is informed, and does not become a second, slower reporting route. |
| The vendor ships a version change | Model owner confirms within ten working days whether the local performance check still holds. If it cannot be confirmed, the tool goes back a tier. |
| A regulator or the vendor issues a safety communication | Chair may suspend without a meeting, and reports the suspension afterwards. |
| Twelve months since the last review, whatever else has happened | Re-approval or retirement. Nothing stays live by inertia. |
The intake form
Fourteen fields, on one page, completed before an item reaches an agenda. The form is the committee’s real gate: most of what should be turned away can be turned away on an incomplete intake, without a meeting and without a debate.
The public forms that exist do a different job. The Health Sector Coordinating Council’s appendix F scores a use case for risk and its appendix G assesses the vendor [23]; CHAI’s playbooks describe how to design an intake process rather than supplying the form [21]. Neither is the short internal form a clinician or a department head fills in to ask for a tool in the first place. That is what this one is.
01 · Requesting service and owner
A named individual who will still be here in a year, and their service line.
02 · The problem
What is being decided badly today, by whom, how often, and what it costs. One paragraph, no product name in it.
03 · Proposed tool and vendor
Product name, version, and whether it sits inside an existing contract.
04 · Intended use, in our words
Which patients, which clinicians, at what point in the pathway, and what the output changes.
05 · Indications for use, in the vendor’s words
The cleared or labelled statement, quoted verbatim. Attach the source document.
06 · Regulatory status
Pathway and submission number, or an explicit statement that the tool is not a regulated device and why.
07 · Change control
Is there a Predetermined Change Control Plan or equivalent, and what does it permit the vendor to change unannounced.
08 · Evidence supplied
External validation sites, populations, years, operating point, calibration, subgroup breakdowns. Attach; do not summarise.
09 · Local validation plan
Which of your own data, how much of it, which metrics, who runs it, and the pass mark — agreed before the result is known.
10 · Population and equity
Subgroups in the affected population, and which of them the vendor’s evidence covers. Name the ones it does not.
11 · Data flow
What leaves the environment, where it is processed, retention, and whether any of it trains a model. Attach the relevant contract clause.
12 · Workflow and burden
Expected alerts or outputs per unit of clinical activity, who absorbs them, and what is being removed to make room.
13 · Monitoring and thresholds
Metrics, cadence, the named owner, and the numeric thresholds that will force this back onto the agenda.
14 · Stop conditions
What would make you turn it off, decided now. If nothing would, the submission is not ready.
The build sequence behind this charter — how the committee is stood up, and which published framework grounds each step — is set out in the hospital AI governance committee playbook.
Artifact 05
What this kit cannot tell you
The failure mode of every resource in this space is false confidence. These are the questions a reader will reasonably have that the published record does not answer, listed so that nobody mistakes the silence for agreement.
Open question · 1
Do any of these frameworks predict whether a tool will work in your service?
No. None has been validated against deployment outcomes. What is documented is that they disagree with each other: a systematic review applied both CLAIM and FUTURE-AI to the same 325 studies of AI in radiological imaging of soft-tissue and bone tumours, published before 17 July 2024, and reported mean scores of 28.9 ± 7.5 out of 53 on CLAIM and 5.1 ± 2.1 out of 30 on FUTURE-AI. The same literature, two different verdicts, because the two instruments are measuring different objects. [1][6][11]
Open question · 2
Which evaluation tool is the authoritative one?
There is not one. A scoping review searching MEDLINE, Embase, CINAHL, PsycINFO and IEEE to April 2024 identified 46 tools — 26 reporting guides, 16 critical appraisal tools, 2 study-quality tools and 2 risk-of-bias tools. Of the frameworks reconciled above, only PROBAST+AI’s own stakeholder table mentions a healthcare professional verifying a model “before purchasing or using” it, and that is one line rather than a procurement method. [2][10]
Open question · 3
What level of performance is good enough to deploy?
No framework sets a number, and none could. The answer depends on the prevalence in your population, on what you currently do instead, and on how asymmetric the harms of a false positive and a false negative are in that specific pathway. Anyone offering a universal threshold is guessing, and the guess is the part you would be buying.
Open question · 4
How many of the AI-enabled devices on the FDA’s list are actually in clinical use?
Unknown, and the list cannot answer it. It counts marketing authorizations, one row per submission number: 1,524 rows, of which 1,466 are 510(k)s, 39 De Novo and 19 PMA — derived by counting the FDA’s own downloadable CSV on 13 August 2026, because the FDA prints no total. Content on that page was current as of 16 June 2026 and the newest authorization on it is dated 30 March 2026. The FDA states plainly that the list “is not a comprehensive resource of AI-enabled medical devices”. Authorization is not adoption, and nothing published converts one into the other. [12]
Open question · 5
If a device is cleared, has anyone tested it on patients like yours?
Usually not, and clearance does not claim it. The FDA describes a 510(k) as a submission demonstrating that a device is substantially equivalent to a legally marketed device, and states that only a small percentage of 510(k)s require clinical data to support clearance. Local validation is treated as the buyer’s job: the FDA’s own draft guidance of 7 January 2025 recommends manufacturers document for users “how to conduct local site-specific acceptance testing or validation”. That guidance was still a draft, and still non-binding, when this page was last checked. [12][14][15]
Open question · 6
Is a hospital AI governance committee actually required?
No. Not by any hospital accreditation standard, and not by regulation. The Joint Commission and the Coalition for Health AI published guidance on 17 September 2025 that names seven elements of responsible AI use, and it describes itself as an initial, high-level document. The Joint Commission’s Responsible Use of AI in Healthcare certification, announced 1 June 2026, is voluntary and separate from accreditation — organisations do not need to be accredited to apply for it. Every governance artifact cited on this page carries its own disclaimer that adoption is voluntary; the CHAI playbooks state that they do not establish or reflect a legal, regulatory, professional or industry standard of care. [20][21][22]
Open question · 7
How many hospitals have one, and does having one help?
The first half has a number and the second half does not. Asked “Who in your hospital or health care system is accountable for evaluating models?”, 66% of the 1,587 US non-federal acute care hospitals that use predictive AI selected a specific committee or task force for machine learning or predictive modelling — from the 2024 American Hospital Association IT Supplement, fielded April to September 2024, response rate 51%. That is a figure about accountability for evaluation among hospitals already using predictive AI, not a count of hospitals with an AI governance committee, and it should not be quoted as one. No study has tested whether having such a committee changes patient outcomes. [30]
Open question · 8
What should the committee be made of?
Nobody has established this, and no survey reports it. The published guidance names types of expertise and stops; the templates released in 2026 leave membership, quorum and decision rights as fields for the adopting organisation to fill in. The descriptions of committees that genuinely operate are single-institution case reports — a comprehensive cancer center published one year of its committee’s work in 2025, covering 26 AI models including large language models, 2 ambient AI pilots and 33 nomograms, and stated that before it, no effective governance model had been reported in oncology. The charter above is a filled-in position offered so there is something specific to disagree with. It is not evidence. [20][21][23][29]
Open question · 9
What do vendors disclose without being asked?
Less than this checklist asks for. Of the 168 machine-learning-enabled Class II devices the FDA authorized in 2024, 49 (29.2%) reported both sensitivity and specificity, 15.5% provided demographic data, and 16.7% of summaries mentioned a Predetermined Change Control Plan. Radiology accounted for 74.4% of that year’s authorizations. The gap between what a checklist requires and what a submission contains is the reason the vendor question list exists. [37]
Artifact 06
Sources
Every one of these is a primary source: the journal article, the regulator's own document, or the standards body's own text. Each was resolved and read on 13 August 2026. Where a figure appears above, its unit, denominator, population and as-of date travel with it.
Lekadir K, Frangi AF, Porras AR, et al. FUTURE-AI: international consensus guideline for trustworthy and deployable artificial intelligence in healthcare. BMJ. 2025;388:e081554. Published 5 February 2025. Six guiding principles and 30 recommendations, from a consortium of 117 experts across 50 countries.
https://doi.org/10.1136/bmj-2024-081554Moons KGM, Damen JAA, Kaul T, et al. PROBAST+AI: an updated quality, risk of bias, and applicability assessment tool for prediction models using regression or artificial intelligence methods. BMJ. 2025;388:e082505. Published 24 March 2025. 34 signalling questions — 16 for model development, 18 for model evaluation — across four domains.
https://doi.org/10.1136/bmj-2024-082505Kwong JCC, Khondker A, Lajkosz K, et al. APPRAISE-AI Tool for Quantitative Evaluation of AI Studies for Clinical Decision Support. JAMA Network Open. 2023;6(9):e2335377. 24 items in six domains, scored out of 100 points.
https://doi.org/10.1001/jamanetworkopen.2023.35377Collins GS, Moons KGM, Dhiman P, et al. TRIPOD+AI statement: updated guidance for reporting clinical prediction models that use regression or machine learning methods. BMJ. 2024;385:e078378. 27 items.
https://doi.org/10.1136/bmj-2023-078378Vasey B, Nagendran M, Campbell B, et al. Reporting guideline for the early stage clinical evaluation of decision support systems driven by artificial intelligence: DECIDE-AI. BMJ. 2022;377:e070904. 17 AI-specific items containing 28 subitems, plus 10 generic items.
https://doi.org/10.1136/bmj-2022-070904Tejani AS, Klontzas ME, Gatti AA, et al. Checklist for Artificial Intelligence in Medical Imaging (CLAIM): 2024 Update. Radiology: Artificial Intelligence. 2024;6(4):e240300. 44 items; the 2020 original (Mongan J, Moy L, Kahn CE Jr. Radiol Artif Intell. 2020;2(2):e200029) had 42. Neither paper states its own total in a sentence.
https://doi.org/10.1148/ryai.240300Liu X, Cruz Rivera S, Moher D, et al. Reporting guidelines for clinical trial reports for interventions involving artificial intelligence: the CONSORT-AI extension. Nature Medicine. 2020;26:1364-1374. 14 new AI-specific items, added to the core CONSORT 2010 checklist.
https://doi.org/10.1038/s41591-020-1034-xNorgeot B, Quer G, Beaulieu-Jones BK, et al. Minimum information about clinical artificial intelligence modeling: the MI-CLAIM checklist. Nature Medicine. 2020;26(9):1320-1324. Six parts; the paper states no item total.
https://doi.org/10.1038/s41591-020-1041-yHernandez-Boussard T, Bozkurt S, Ioannidis JPA, Shah NH. MINIMAR (MINimum Information for Medical AI Reporting): Developing reporting standards for artificial intelligence in health care. Journal of the American Medical Informatics Association. 2020;27(12):2011-2015. Four components, 21 named reporting fields in the authors’ own table; published explicitly as a starting point for discussion.
https://doi.org/10.1093/jamia/ocaa088Cabello JB, Ruiz Garcia V, Torralba M, et al. Critical Appraisal Tools for Evaluating Artificial Intelligence in Clinical Studies: Scoping Review. Journal of Medical Internet Research. 2025;27:e77110. MEDLINE, Embase, CINAHL, PsycINFO and IEEE searched to April 2024; 46 tools identified — 26 reporting guides, 16 critical appraisal tools, 2 study-quality tools, 2 risk-of-bias tools.
https://doi.org/10.2196/77110Spaanderman DJ, Marzetti M, Wan X, et al. AI in radiological imaging of soft-tissue and bone tumours: a systematic review evaluating against CLAIM and FUTURE-AI guidelines. eBioMedicine. 2025;114:105642. 15,015 abstracts screened, 325 articles included, published before 17 July 2024; mean CLAIM score 28.9 ± 7.5 of 53, mean FUTURE-AI score 5.1 ± 2.1 of 30.
https://doi.org/10.1016/j.ebiom.2025.105642US Food and Drug Administration. Artificial Intelligence-Enabled Medical Devices (device list and downloadable CSV). Content current as of 16 June 2026; newest authorization on the list dated 30 March 2026. Counts on this page were derived by parsing the FDA’s own CSV on 13 August 2026 — the FDA publishes no total, and states that the list “is not a comprehensive resource of AI-enabled medical devices”.
https://www.fda.gov/medical-devices/software-medical-device-samd/artificial-intelligence-enabled-medical-devicesUS Food and Drug Administration. Marketing Submission Recommendations for a Predetermined Change Control Plan for Artificial Intelligence-Enabled Device Software Functions — Guidance for Industry and Food and Drug Administration Staff. Final guidance, docket FDA-2022-D-2628; originally issued 4 December 2024, reissued 18 August 2025. Names three required sections: Description of Modifications, Modification Protocol, and Impact Assessment.
https://www.fda.gov/regulatory-information/search-fda-guidance-documents/marketing-submission-recommendations-predetermined-change-control-plan-artificial-intelligenceUS Food and Drug Administration. Premarket Notification 510(k). Content current as of 22 August 2024. A 510(k) demonstrates that a device is substantially equivalent to a legally marketed device; the FDA states that only a small percentage of 510(k)s require clinical data to support clearance.
https://www.fda.gov/medical-devices/premarket-submissions-selecting-and-preparing-correct-submission/premarket-notification-510kUS Food and Drug Administration. Artificial Intelligence-Enabled Device Software Functions: Lifecycle Management and Marketing Submission Recommendations. DRAFT guidance, docket FDA-2024-D-4488, issued 7 January 2025. Recommends manufacturers document for users “how to conduct local site-specific acceptance testing or validation”. Confirmed still draft and non-binding on 13 August 2026.
https://www.fda.gov/regulatory-information/search-fda-guidance-documents/artificial-intelligence-enabled-device-software-functions-lifecycle-management-and-marketingHealth Data, Technology, and Interoperability: Certification Program Updates, Algorithm Transparency, and Information Sharing (HTI-1), final rule. 89 Federal Register 1192, 9 January 2024; effective 8 February 2024. Decision Support Intervention criterion codified at 45 CFR 170.315(b)(11): 13 source attributes for evidence-based DSIs and 31 for Predictive DSIs across nine categories, with a requirement to indicate when an attribute is not available for review.
https://www.federalregister.gov/documents/2024/01/09/2023-28857/health-data-technology-and-interoperability-certification-program-updates-algorithm-transparencyRegulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act). Official Journal of the European Union, 12 July 2024. Article 6(1) and Annex I; Article 14; Article 50(4); Article 113.
https://eur-lex.europa.eu/eli/reg/2024/1689/ojRegulation (EU) 2026/1744 of 8 July 2026 amending Regulation (EU) 2024/1689 as regards the simplification of the implementation of harmonised rules on artificial intelligence. Official Journal, 24 July 2026; in force 27 July 2026. Amends Article 113 so high-risk obligations apply from 2 December 2027 for Annex III systems and 2 August 2028 for Annex I systems, the route medical devices take.
https://eur-lex.europa.eu/eli/reg/2026/1744/ojRegulation (EU) 2017/745 on medical devices (MDR), Annex VIII, Chapter III, rule 11. Official Journal, 5 May 2017. Software intended to provide information used to take decisions with diagnosis or therapeutic purposes is class IIa, rising to IIb or III with the severity of the decision it influences.
https://eur-lex.europa.eu/eli/reg/2017/745/ojJoint Commission and Coalition for Health AI. Guidance on the Responsible Use of AI in Healthcare (RUAIH). 17 September 2025, 8 pages. Names seven elements of responsible AI use and a list of expertise a governance structure should span; describes itself as “an initial, high-level document”. Voluntary.
https://assets.ctfassets.net/7s4afyr9pmov/5RX4XUbRg0l1JG0gwXGkTM/bea3c1949d6d8cdd8ec2ae2f82f737ad/JC-CHAI_RUAIH_Guidance.pdfCoalition for Health AI. Responsible AI Governance Playbooks, v1.0. Released 27 May 2026. Eight playbooks; domain 2 covers organizational structures, with four baseline controls (critical roles, a formal AI governance committee, an intake and monitoring process, escalation protocols), a worked three-tier escalation example, and case studies from Johns Hopkins, Mayo Clinic and TrueCare. Its own disclaimer states it “does not establish, define, or reflect a legal, regulatory, professional, or industry standard of care”.
https://www.chai.org/news/coalition-for-health-ai-chai-releases-comprehensive-governance-playbooks-toJoint Commission. Responsible Use of AI in Healthcare (RUAIH) certification, announced 1 June 2026. Voluntary and separate from accreditation: organizations “do not need to be accredited by Joint Commission to apply”, and the programme does not certify individual AI products.
https://www.jointcommission.org/en-us/knowledge-library/news/2026-05-responsible-use-of-ai-in-healthcare-certificationHealth Sector Coordinating Council Cybersecurity Working Group. Health Industry AI Cyber Governance Framework Implementation Guide. May 2026, 87 pages. Contains an AI governance committee charter template (appendix J), a RACI matrix (appendix B), and a use-case justification and risk-scoring form (appendix F). Explicitly scoped to the cybersecurity dimensions of AI governance, and explicitly not required.
https://healthsectorcouncil.org/wp-content/uploads/2026/05/AI-Cyber-Governance-Framework-Implementation-Guide.pdfNational Academy of Medicine. An Artificial Intelligence Code of Conduct for Health and Medicine: Essential Guidance for Aligned Action. National Academies Press, 14 July 2025. Ten Code Principles and six Code Commitments. The full text contains no guidance on governance committees or intake.
https://doi.org/10.17226/29087Health AI Partnership (Duke Institute for Health Innovation). Key decisions in adopting an AI solution — eight key decision points across four phases (procurement, development, integration, lifecycle management), with 31 best-practice guides. Accessed 13 August 2026.
https://healthaipartnership.org/key-decisions-in-adopting-an-ai-solutionAmerican Medical Association, with Manatt Health. Governance for Augmented Intelligence. AMA STEPS Forward toolkit, published 8 May 2025. Eight steps, the fifth of which is to define project intake, vendor evaluation and assessment processes.
https://edhub.ama-assn.org/steps-forward/module/2833560World Health Organization. Ethics and governance of artificial intelligence for health. 28 June 2021, 150 pages, ISBN 978-92-4-002920-0. Six core principles. Companion guidance on large multi-modal models published 2024, ISBN 978-92-4-008475-9.
https://www.who.int/publications/i/item/9789240029200Bedoya AD, Economou-Zavlanos NJ, Goldstein BA, et al. A framework for the oversight and local deployment of safe and high-quality prediction models. Journal of the American Medical Informatics Association. 2022;29(9):1631-1636. A published oversight framework from one academic health system.
https://doi.org/10.1093/jamia/ocac078Stetson PD, Choy J, Summerville N, et al. Responsible Artificial Intelligence governance in oncology. npj Digital Medicine. 2025;8(1):407. One year of a comprehensive cancer center’s AI Governance Committee: registration and monitoring of 26 AI models including large language models, 2 ambient AI pilots, and review of 33 nomograms. The authors state that before this report, no effective governance models had been reported in oncology. Retrieved via PubMed (PMID 40615544).
https://doi.org/10.1038/s41746-025-01794-wChang W, Owusu-Mensah P, Everson J, Richwine C. Hospital Trends in the Use, Evaluation, and Governance of Predictive AI, 2023-2024. ASTP/ONC Data Brief No. 80, September 2025. Based on the 2024 American Hospital Association Information Technology Supplement, fielded April to September 2024, response rate 51%, N = 2,253 non-federal acute care hospitals. Asked “Who in your hospital or health care system is accountable for evaluating models? (Check all that apply.)”, 66% of the 1,587 hospitals that use predictive AI selected “Specific Committee or Task Force for Machine Learning or Predictive Modeling”.
https://www.healthit.gov/data/data-briefs/hospital-trends-use-evaluation-and-governance-predictive-ai-2023-2024Kim DW, Jang HY, Kim KW, Shin Y, Park SH. Design Characteristics of Studies Reporting the Performance of Artificial Intelligence Algorithms for Diagnostic Analysis of Medical Images. Korean Journal of Radiology. 2019;20(3):405-410. Of 516 eligible studies, 31 (6%) performed external validation.
https://doi.org/10.3348/kjr.2019.0025Wong A, Otles E, Donnelly JP, et al. External Validation of a Widely Implemented Proprietary Sepsis Prediction Model in Hospitalized Patients. JAMA Internal Medicine. 2021;181(8):1065-1070. External AUC 0.63; at the live alerting threshold the model did not identify 1,709 of 2,552 patients with sepsis (67%).
https://doi.org/10.1001/jamainternmed.2021.2626Van Calster B, McLernon DJ, van Smeden M, Wynants L, Steyerberg EW. Calibration: the Achilles heel of predictive analytics. BMC Medicine. 2019;17:230.
https://doi.org/10.1186/s12916-019-1466-7Nagendran M, Chen Y, Lovejoy CA, et al. Artificial intelligence versus clinicians: systematic review of design, reporting standards, and claims of deep learning studies in medical imaging. BMJ. 2020;368:m689. 61 of 81 non-randomised studies claimed performance comparable to or better than clinicians.
https://doi.org/10.1136/bmj.m689Andaur Navarro CL, Damen JAA, Takada T, et al. Risk of bias in studies on prediction models developed using supervised machine learning techniques: systematic review. BMJ. 2021;375:n2281. 148 of 152 studies (87%, 95% CI 81% to 91%) were at high risk of bias; 56% had an inadequate number of events per candidate predictor; 39% assessed overfitting improperly.
https://doi.org/10.1136/bmj.n2281Obermeyer Z, Powers B, Vogeli C, Mullainathan S. Dissecting racial bias in an algorithm used to manage the health of populations. Science. 2019;366(6464):447-453.
https://doi.org/10.1126/science.aax2342Almarie B, Gonzalez-Gonzalez LF, dos Santos Barbosa LA, et al. Machine Learning-Enabled Medical Devices Authorized by the US Food and Drug Administration in 2024: Regulatory Characteristics, Predicate Lineage, and Transparency Reporting. Biomedicines. 2025;13(12):3005. Of 168 ML-enabled Class II devices authorized in 2024, 49 (29.2%) reported both sensitivity and specificity, 15.5% provided demographic data, and 16.7% of summaries mentioned a Predetermined Change Control Plan; radiology accounted for 74.4%. Retrieved via PubMed.
https://doi.org/10.3390/biomedicines13123005Davis SE, Lasko TA, Chen G, Siew ED, Matheny ME. Calibration drift in regression and machine learning models for acute kidney injury. Journal of the American Medical Informatics Association. 2017;24(6):1052-1061. Discrimination was maintained across nine years of validation while calibration declined and all models increasingly over-predicted risk.
https://doi.org/10.1093/jamia/ocx030American Medical Association, Center for Digital Health and AI. 2026 Physician Survey on Augmented Intelligence, March 2026. 1,692 US physicians participated; fielded 15 January to 2 February 2026. 81% reported any awareness or use of AI (n = 1,342 for that item). The AMA notes the 2026 wave includes qualified partial responses where earlier waves did not, so it is not a like-for-like comparison with 2023 or 2024.
https://www.ama-assn.org/system/files/physician-ai-sentiment-report.pdf
Keep it current
The kit is free. Joining is how you stay current with it.
Regulators move and frameworks get revised. Members get this page's revisions, the whole library of 104 pieces, the paper feed, and one paper a week read closely by email.
Free · one email address