Most explanations of the UK's approach to AI in healthcare stop at a sentence: "the MHRA runs a regulatory sandbox called the AI Airlock." That is true and almost useless. The value of the Airlock is in what it actually tested and what it found — the concrete regulatory gaps that only appear when a real AI product meets the real rulebook. This page walks through both of the MHRA's published cohorts, mapping each gap to the case study that surfaced it, using the pilot and Phase 2 reports themselves as the source. It is an explainer for orientation and general information; confirm the current framework and its application to any specific product with regulatory counsel before acting. As of July 2026.
What the AI Airlock is
The AI Airlock is the MHRA's first regulatory sandbox for AI as a medical device (AIaMD), developed and delivered with the NHS and the Department of Health and Social Care (DHSC) 3. The MHRA announced it in May 2024 to test real-world products and prototypes of AI medical devices, working with a coalition of UK Approved Bodies, the NHS AI Lab and DHSC 2. AIaMD is a subset of software as a medical device, regulated in the UK under the Medical Devices Regulations 2002.
The purpose is narrow and deliberate: to identify and address the challenges of regulating standalone AI intended for direct clinical use, so that guidance and rules can evolve on the strength of real evidence rather than speculation. A regulatory sandbox does this by taking the regulator into uncertain areas that "often cannot be explored within normal business processes" 3. The Airlock is a learning and testing mechanism rather than a route to market — a distinction worth holding on to, because its most useful outputs are the gaps it exposes.
The pilot built three testing environments, and later phases reuse them 3:
- Simulation Airlock — a structured workshop or roundtable that runs a "pre-mortem" on a challenge with regulators, clinicians, technical experts, academics and legal specialists, for problems at concept or early-development stage.
- Virtual / Research Airlock — a "test bed" where an AI model is exercised on real or synthetic data so testers can probe how it responds and how its internal process works, before any real-world deployment.
- Real-world Airlock — deployment in a hospital or trust alongside clinicians and anonymised patient data, where the AI is observed in workflow but its outputs are not used to make decisions about patients. This environment is run to protect clinical pathways and is distinct from a clinical investigation.
The pilot: four case studies, four gaps
The pilot ran from April 2024 to March 2025. A public call drew 40 applications — 31% using generative AI, 23% machine learning, 13% predictive AI and 10% multimodal systems, with diagnostics (25%), screening or imaging (19%) and clinical decision support (19%) the most common uses. Five candidates were taken forward and four completed the full pilot 3. Each was chosen because it pressed on a distinct, unresolved regulatory question.
| Case study | Regulatory gap it probed | What the Airlock found |
|---|---|---|
| Synthetic data for a radiology-report tool | Whether synthetic training/validation data can evidence compliance | 4,000 synthetic reports were generated by an LLM; using one LLM to judge another raised a "circularity" risk that outputs could reinforce shared errors 3 |
| A guideline-grounded clinical LLM | How to manage LLM hallucination and non-determinism | Retrieval grounding in trusted guidelines cut hallucinations to zero across 436 questions, versus 23 for a baseline, with consistent repeat answers 3 |
| An oncology decision-support tool | The trade-off between explainability and clinical performance | Multiple explainability methods showed promise, but stakeholders agreed their use needs clearer regulatory direction across different audiences 3 |
| A real-time monitoring platform | Continuous post-market surveillance of a live AIaMD | Across 180 reports at two hospital sites, real-time monitoring surfaced data-quality issues, model drift and a pattern of clinician over-reliance 3 |
Two of these findings are worth drawing out, because they turn abstract worries into numbers.
Hallucination is tractable, and the Airlock measured it. The pilot worked with a clinical large language model grounded in verified sources — NICE guidelines — using retrieval-augmented generation. In an experiment of 436 clinical questions, the retrieval-enabled system produced no hallucinations, against 23 from a baseline model, and returned highly similar answers to repeated prompts 3. On that evidence the MHRA's programme recommended that guidance should encourage safety measures such as retrieval grounding, "given the demonstrated ability to reduce hallucinations and support safer, more explainable outputs" 3. For the failure mode this addresses, see our glossary entry on AI hallucination in clinical contexts.
Over-reliance is a real-world risk, and monitoring can see it. The real-time surveillance case study reviewed 180 reports across two hospital sites and used continuous monitoring to detect data-quality issues, model drift and automation bias. At one site it identified a pattern of over-reliance on the AIaMD, with contributing factors including clinician fatigue, time pressure and experience level 3. This is the argument for treating deployment as an ongoing surveillance problem — the discipline our glossary calls algorithmovigilance — rather than a one-time clearance.
A third case study is worth a moment because it names a trap specific to generative AI. Working on a tool that drafts the "impression" section of radiology reports, the Airlock examined the generation and validation of 4,000 synthetic reports produced by an LLM. The catch it surfaced: when one LLM is used to judge the outputs of another, errors and biases can be reinforced rather than independently caught — a "circularity" risk that leaves the evidence less trustworthy than it looks. Current device regulation offers no settled guidance on validating text-based synthetic data, and the Airlock's recommendation was that the MHRA develop best practice for generating, validating and assessing it, particularly where synthetic data supports a regulatory submission 3.
Across all four projects the recurring themes were risk management, validation of text-based data from large language models, AI errors and non-determinism, explainability, and post-market surveillance 3. The pilot report is explicit about its own standing: "This report does not constitute formal MHRA guidance" 3. Its role is to feed the National AI Commission and the MHRA's Software and AI as a Medical Device Change Programme.
Phase 2: three unresolved challenges
Phase 2 ran from April 2025 to March 2026, building directly on the pilot. It drew 51 applications and selected a cohort of seven candidates, supported by DHSC and the Regulatory Innovation Office, and was structured around the three challenges the pilot flagged as critical and unresolved 4:
| Challenge area | The question it asks |
|---|---|
| Scope of intended purpose and validation | How to define, maintain and enforce the intended purpose of AI that can evolve in function and clinical influence over time 4 |
| Predetermined change control plans and post-market surveillance | How to manage iterative updates to AI safely and proportionately across the device lifecycle 4 |
| Performance evaluation for AI-powered in vitro diagnostics | How to assess and evidence the performance of AI-powered IVDs 4 |
The second of these connects directly to a mechanism worth knowing on its own terms: the predetermined change control plan, which lets a manufacturer pre-specify how a model may change without a fresh submission each time. Phase 2's recommendation was to develop principles-based PCCP guidance — clarity on the limits of pre-authorised change, and on when a change to a device is significant enough to fall outside the plan 4.
Phase 2 also produced a sharper set of insights than the pilot. Three stand out 4. First, "statistically significant change and clinically important change are not the same thing," so acceptance criteria and thresholds should be grounded in clinical relevance rather than statistics alone. Second, human oversight may degrade over a device's life: as users grow familiar and the system stays accurate, they "may trend toward applying less scrutiny," which weakens human-in-the-loop safeguards over time. Third, LLM-based products without active guardrails "may begin to exhibit functions beyond their intended purpose" — the scope-creep problem that makes intended-purpose control hard for generative systems.
Phase 2 was independently evaluated by a research agency, which reported strong stakeholder satisfaction and a consistent call for the programme's insights to be translated into clearer, system-level guidance 4. As with the pilot, the Phase 2 materials carry a disclaimer that the innovator case studies "do not constitute MHRA guidance or policy" 4.
What comes next
A Phase 3, funded by DHSC for three years from April 2026, will focus on turning the programme's evidence into tangible MHRA actions and policy, shaped by the recommendations of the National Commission into the Regulation of AI in Healthcare expected during 2026 4. The published record so far sits on the MHRA's AI Airlock collection page: the pilot report (October 2025), the Phase 2 report (published 9 June 2026, updated 17 July 2026), and the associated cohort and workshop material 15.
How to read this
Four cautions travel with everything above.
First, the Airlock produces evidence and recommendations, not law. Both reports state plainly that they do not constitute formal MHRA guidance 34. The findings shape where the rules are heading; they do not change a manufacturer's current obligations.
Second, treat the case-study results as illustrative rather than generalisable. Zero hallucinations across 436 questions in one grounded system, or an over-reliance pattern at one of two sites, are pointed demonstrations from small, bespoke tests — signals of what is possible and what to watch, rather than population-level performance claims. The discipline for reading any such number lives in our companion material on evaluation.
Third, the sandbox is a UK device-framework instrument, and healthcare AI is governed differently elsewhere. The Airlock works within the Medical Devices Regulations 2002; a product placed on the EU market faces the EU AI Act's medical-device route instead, and the wider cross-jurisdiction picture lives in our global AI in health regulation tracker.
Fourth, status is perishable. Cohorts close, reports publish, and a new phase is already funded — which is why this page carries a ninety-day cadence and a date on its face. Because the subject is device regulation, treat this explainer as orientation only and confirm the current requirements for any specific product with your regulatory or compliance function before you rely on them.
Sources and method
This explainer is built from primary MHRA sources: the AI Airlock collection page on GOV.UK 1, the May 2024 launch press release 2, the pilot programme report of October 2025 3, and the Phase 2 programme report of June 2026 with its publication record 45. Every figure — the 40 and 51 applications, the four and seven completed candidates, the 436-question hallucination result, the 180-report monitoring study — is drawn from the report cited beside it. We revisit this page every ninety days and whenever the MHRA opens a new cohort or publishes new Airlock outputs. Dates and findings are current as of July 2026.