Regulation

WHO guidance on large language models in health

What the World Health Organization's 2024 guidance on large multi-modal models actually says — the five ways it expects these systems to be used in health, the risks it names, who it tells to do what, and the one thing to keep straight: it is advisory, and binding rules sit elsewhere. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • The WHO released 'Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models' on 18 January 2024, with more than 40 recommendations for governments, technology companies, and health-care providers.
  • Large multi-modal models (LMMs) accept more than one type of input — text, images, video — and generate varied outputs; the guidance covers the generative systems clinicians now actually encounter.
  • It maps five application areas: diagnosis and clinical care; patient-guided use; clerical and administrative tasks; professional education; and scientific research and drug development.
  • It names concrete risks: false, inaccurate, or biased outputs; automation bias, where errors are overlooked that would otherwise be caught; biased training data; and cybersecurity exposure of patient information.
  • The single most important thing to keep straight: WHO guidance is advisory. Binding market-authorisation rules sit with device regulators — the EU AI Act, the UK MHRA, and national authorities — not with the WHO.

When people say "the WHO has guidance on AI chatbots in healthcare," they are usually pointing at one document: Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models, released on 18 January 2024 1. It is the most widely cited global statement on generative AI in health, and it is also widely misread — treated as either a binding rulebook it never was, or a vague ethics essay it is not. This page sets out what the guidance actually says: the systems it covers, the five ways it expects them to be used, the risks it names, who it tells to act, and where it sits relative to the binding rules that decide whether a tool can reach patients. As of July 2026.

What is the WHO guidance?

The 2024 guidance carries more than 40 recommendations for governments, technology companies, and health-care providers 2. It builds on the WHO's 2021 foundation, Ethics and governance of artificial intelligence for health, which set out six consensus principles for AI in health 3 — reproduced at the end of this page, because the 2024 recommendations sit on top of them rather than replacing them.

The reason a fresh document was needed is in its title. The 2021 guidance predated the public arrival of general-purpose generative systems; the 2024 guidance is the WHO's response to them. It focuses specifically on large multi-modal models (LMMs) — which the WHO describes as able to "accept one or more type of data inputs, such as text, videos, and images, and generate diverse outputs," and as "unique in their mimicry of human communication and ability to carry out tasks they were not explicitly programmed to perform" 2. An LMM is a superset of the text-only clinical large language model; both are built on the same underlying foundation models. The multi-modal framing is deliberate and matters in health, where a single question can span an image, a lab value, and a free-text history at once. A system that can take in more than one of those and answer in plain language is exactly what the guidance is written for — and exactly what is now reaching clinics, and patients, directly.

What are the five application areas?

The guidance does not treat "AI in health" as one undifferentiated thing. It maps five distinct areas where LMMs are being applied, and the distinction matters because the risks and the appropriate safeguards differ sharply across them 2.

#Application areaWhat it looks like in practice
1Diagnosis and clinical careAssisting with clinical decisions, interpretation, and management
2Patient-guided useA patient querying symptoms or treatment options directly
3Clerical and administrative tasksDocumentation, cataloguing, and record-keeping in electronic records
4Professional educationTraining of clinicians and nurses
5Scientific research and drug developmentLiterature synthesis, discovery, and research support

The split is useful on its own. A model drafting an administrative note (area 3) and a model a patient consults unsupervised about symptoms (area 2) raise different orders of risk, even if it is literally the same underlying system. The guidance's central move is to insist that governance be matched to the use, not to the technology in the abstract.

Which risks does the WHO name?

The guidance is specific about what can go wrong, and its list maps closely onto the failure modes practitioners already worry about 2:

  • False, inaccurate, or biased outputs — statements that "could harm people using such information in making health decisions." This is the clinical face of what our glossary calls AI hallucination in clinical contexts.
  • Automation bias — the risk that with an automated system in the loop, "errors are overlooked that would otherwise have been identified." This is the specific reason a nominal human-in-the-loop safeguard can quietly weaken in practice.
  • Biased training data — data "of poor quality or biased, whether by race, ethnicity, ancestry, sex, gender identity, or age," reproduced downstream.
  • Cybersecurity and data exposure — vulnerabilities that endanger patient information.

Naming these is the guidance's most practically useful contribution: it gives a health system a shared vocabulary for the specific hazards to test for before deployment, rather than a general appeal to be careful.

The weight of these risks is spread unevenly across the five areas, which is the reason the guidance separates them in the first place. A model drafting an administrative note sits behind a clinician who will read what it produced; a model a patient consults directly about symptoms has no such backstop, so patient-guided use concentrates the danger of a confident, wrong answer. Diagnosis and clinical care raise the automation-bias problem most sharply, because the entire purpose of the tool is to influence a decision a human then rubber-stamps. Matching the depth of governance to the exposure of the use — heavier for area 1 and 2, lighter for a supervised area 3 — is the guidance's organising idea, and a more useful instruction than a single standard applied identically everywhere 2.

Who does the WHO tell to act?

The recommendations are addressed to distinct actors, and the division of labour is the part most summaries drop 2.

Governments are asked to invest in or provide public, not-for-profit infrastructure including computing power; to use laws, policies, and regulations so that LMMs used in health meet ethical obligations and human-rights standards; to assign an existing or new regulatory agency to assess and approve LMMs intended for health use; and to introduce mandatory post-release auditing and impact assessments by independent third parties.

Developers are asked to design LMMs with users and stakeholders — clinicians, researchers, health-care professionals, and patients — involved "from the earliest stages," rather than by scientists and engineers alone, and to design systems "to perform well-defined tasks with the necessary accuracy and reliability," with an ability to predict secondary outcomes.

Health-care providers, the third audience, are the deployers who must weigh those tools in real settings — the group for whom the risk list above is a working checklist. They sit closest to the failure modes, too: the clinician who accepts a fabricated line in a drafted note, or the triage step that inherits a model's blind spot. The guidance's implicit message to them is that adopting an LMM is a clinical-governance decision, subject to the same scrutiny as any other change to how care is delivered, rather than an IT purchase to be waved through.

The point everyone gets wrong: advisory, not binding

If you remember one thing about the WHO guidance, make it this. The guidance is advisory. It shapes national policy and sets expectations for responsible design and use, but it does not itself authorise or prohibit a product. The WHO's own division of labour makes this explicit: it recommends that governments assign a regulatory agency to assess and approve health LMMs 2 — an acknowledgement that the WHO is not that agency.

Binding market-authorisation rules live with device regulators. In the EU, AI that meets the medical-device definition is caught by the EU AI Act's staggered timeline layered on existing device law. In the UK, the MHRA regulates AI as a medical device and has been stress-testing the rules through its AI Airlock sandbox. The WHO's own 2023 Regulatory considerations on artificial intelligence for health is itself framed as a high-level resource of concepts and emerging good practice — covering performance evaluation and monitoring, and risk–benefit assessment — for those national regulators, rather than as a rule they must follow 4. Its scope reads as a menu a national authority can draw on when it writes binding rules of its own — documentation and transparency, intended use and validation, data quality, and privacy among the considerations it sets out — which underlines the division of labour: the WHO supplies the thinking, and the regulator supplies the text with legal force 4. For the full cross-jurisdiction map of what is binding where, see our global AI in health regulation tracker.

Read the WHO guidance, then, as the ethical and governance layer that national rules are expected to honour — influential, globally endorsed, and genuinely useful for shaping policy and internal standards — while remembering that the document deciding whether a specific tool can be marketed is a regulator's, not the WHO's.

The 2021 foundation the 2024 guidance builds on

The six principles from the 2021 guidance remain the base layer for everything above 3:

  1. Protect autonomy.
  2. Promote human well-being, human safety, and the public interest.
  3. Ensure transparency, explainability, and intelligibility.
  4. Foster responsibility and accountability.
  5. Ensure inclusiveness and equity.
  6. Promote AI that is responsive and sustainable.

What does it mean for a health system today?

For an organisation deciding how to use these tools now, the guidance converts into a short set of practical moves. Treat the five application areas as distinct governance tracks rather than one blanket policy: the controls that suit an administrative drafting tool are too light for a patient-facing symptom checker. Build the risk list into procurement and evaluation, so a tool is tested for inaccurate output, bias, and data-exposure before it reaches a patient rather than after. Plan for the WHO's call for independent post-release auditing and impact assessments 2 — the monitoring does not stop at go-live. And keep a named human accountable for outputs in the higher-stakes tracks, remembering the automation-bias warning that a nominal reviewer can quietly stop reviewing. None of this waits on a binding rule; it is governance a health system can adopt on its own authority today, and doing so is how the advisory guidance becomes real.

How to read this

Three cautions travel with the summary above.

First, advisory is not the same as optional. The guidance carries weight with national policymakers and funders, and a health system that ignores it is out of step with a globally endorsed standard — even though no one will withhold a CE mark or a clearance purely on WHO grounds. Take it as the reference for internal governance and procurement standards.

Second, the field moves faster than the guidance. The 2024 document was written for the LMMs of its moment; capabilities and deployment patterns keep shifting, and the WHO may issue further guidance. Where the specifics of a model outrun the text, the six principles and the risk list remain the durable part.

Third, guidance is not jurisdiction-specific compliance. For any real deployment, whether a tool needs authorisation, and under which regime, is a question for the relevant regulator and, where liability or data protection is in play, for legal and compliance counsel — confirm it there rather than inferring it from the WHO text.

Sources and method

This page is drawn from primary WHO sources: the 2024 LMM guidance and its accompanying news release 12, the 2021 ethics-and-governance guidance for the six principles 3, and the 2023 regulatory-considerations document for the advisory-versus-binding frame 4. The recommendation figure — more than 40 — the five application areas, the named risks, and the government and developer duties are each taken from the WHO source cited beside them. We revisit this page every ninety days and whenever the WHO issues new or updated AI-for-health guidance. Content is current as of July 2026.

Questions & answers

  • What does the WHO guidance on large multi-modal models say?

    Released on 18 January 2024, it offers more than 40 recommendations on the ethics and governance of large multi-modal models (LMMs) in health. It maps five application areas — diagnosis and clinical care, patient-guided use, administrative tasks, professional education, and research and drug development — names risks such as inaccurate or biased outputs and automation bias, and assigns duties to governments, developers, and health-care providers.

  • Is the WHO guidance legally binding?

    No. It is advisory. It carries real influence over national policy and sets expectations for responsible design and use, but it does not itself grant or deny market authorisation. Binding rules for AI that meets the definition of a medical device sit with device regulators such as the EU under the AI Act and the UK MHRA.

  • What is a large multi-modal model?

    An LMM is a generative model that can accept more than one type of data input — for example text, images, and video — and produce diverse outputs that are not limited to the input type. The WHO uses the term to cover the general-purpose generative systems now reaching health settings, a superset of the text-only large language model.

Sources

  1. World Health Organization. Ethics and governance of artificial intelligence for health: Guidance on large multi-modal models. 18 January 2024. ISBN 978-92-4-008475-9. www.who.int/publications/i/item/9789240084759
  2. World Health Organization. WHO releases AI ethics and governance guidance for large multi-modal models (news release, 18 January 2024). www.who.int/news/item/18-01-2024-who-releases-ai-ethics-and-governance-guidance-for-large-multi-modal-models
  3. World Health Organization. Ethics and governance of artificial intelligence for health: WHO guidance. 28 June 2021. ISBN 978-92-4-002920-0. www.who.int/publications/i/item/9789240029200
  4. World Health Organization. Regulatory considerations on artificial intelligence for health. 19 October 2023. ISBN 978-92-4-007887-1. www.who.int/publications/i/item/9789240078871