Glossary

Foundation model

What a foundation model is, where the term came from, how the idea landed in healthcare — from retinal imaging to pathology — and how the EU AI Act and WHO define and govern the class. As of July 2026.

By Jonas WeirReviewed by Jonas Weir · editorial reviewUpdated

The short version

  • A foundation model is trained on broad data at scale — usually by self-supervision — and adapted to many downstream tasks, instead of being built for one task.
  • The term was coined at Stanford in August 2021; the EU AI Act regulates the same class as 'general-purpose AI models'.
  • Healthcare instantiations are concrete: RETFound learned from 1.6 million unlabelled retinal images; UNI was pretrained on over 100 million pathology image patches from 100,000+ slides.
  • The promise is label efficiency — strong task performance from few labelled examples — and the risk is that one shared base propagates its flaws into every downstream use.

A foundation model is an AI model trained on broad data at scale, usually by self-supervision, that can be adapted to a wide range of downstream tasks 1. In healthcare, one pretrained base can be specialized to read retinal images, classify pathology slides, or draft clinical text — tasks that previously each required their own model.

Why it matters in healthcare

Healthcare AI spent two decades building one narrow model per task, each demanding thousands of expert-labelled examples. Foundation models invert the economics: learning happens once, on massive unlabelled data, and each new task is an adaptation rather than a fresh build. A 2023 Nature perspective formalized where this leads for healthcare — generalist medical AI (GMAI), models able to carry out diverse tasks using very little or no task-specific labelled data 2. The same perspective sketches the reach: one model flexibly interpreting combinations of imaging, electronic health records, laboratory results, genomics, and clinical text, and producing free-text explanations or annotations in return 2. For institutions that could never assemble task-scale labelled datasets, label efficiency is the headline benefit.

How it works in practice

The recipe has two stages. Pretraining: the model learns structure from broad unlabelled data by predicting held-back parts of its own input — self-supervision, no annotators required. Adaptation: the pretrained base is fine-tuned, prompted, or lightly retrained for a specific downstream task using comparatively few labelled examples. A clinical LLM is this pattern applied to language; the same pattern now runs across imaging and signals.

The term itself was coined in the August 2021 Stanford CRFM report, which defined the class by exactly those two properties — training on broad data at scale, adaptability across downstream tasks — and warned that defects in a shared base propagate to everything built on it 1.

Where it appears today

Two healthcare instantiations show the pattern concretely. RETFound (2023) was trained by self-supervised learning on 1.6 million unlabelled retinal images, then adapted with explicit labels; it consistently outperformed comparison models on eye-disease detection and on incident prediction of systemic conditions such as heart failure and myocardial infarction, using fewer labelled examples 3. UNI (2024) did the same for pathology: pretrained on more than 100 million tissue-image patches drawn from over 100,000 diagnostic whole-slide images across 20 tissue types 4.

Regulators now define the class in law. The EU AI Act — Regulation (EU) 2024/1689 — governs "general-purpose AI models": models trained with large amounts of data using self-supervision at scale that display significant generality and competently perform a wide range of distinct tasks 5. The WHO's January 2024 guidance treats large multi-modal models, one prominent foundation-model family, as a distinct governance object for health 6. As of July 2026, both frameworks shape how health systems procure and deploy these models.

Common misunderstandings

Foundation model means language model. Language is one modality. RETFound and UNI are foundation models built purely on images 3 4.

General pretraining guarantees local performance. Adaptation still needs external validation on the deploying population, and shared-base flaws propagate downstream — the Stanford report's central caution 1. Performance can also decay in deployment as populations shift; see model drift.

The regulatory term matches the research term. The EU AI Act says "general-purpose AI model" and attaches specific provider duties to it; the research literature says "foundation model" and means the training paradigm. The overlap is large, and the legal definition is the one that binds 5.

Related terms

Questions & answers

  • Is every large language model a foundation model?

    Large language models are the best-known foundation models, and the class is wider — it includes vision models such as RETFound (retinal images) and UNI (pathology), and multimodal models spanning text, imaging, and signals. The defining traits are broad-data pretraining and adaptability across tasks, whatever the modality.

  • Why do foundation models matter for healthcare specifically?

    Labelled clinical data is scarce and costly. Foundation models front-load learning onto huge unlabelled datasets, so a hospital can adapt one to a local task with far fewer labelled examples — RETFound demonstrated exactly this label efficiency for eye disease and even cardiovascular prediction from retinal images.

  • How does the EU AI Act treat foundation models?

    Under the name "general-purpose AI models" — models trained with large amounts of data using self-supervision at scale that display significant generality. Providers carry documentation and transparency duties, with an added tier for models posing systemic risk, and healthcare deployers inherit obligations when such models power high-risk uses.

Sources

  1. Bommasani R, Hudson DA, Adeli E, et al. On the Opportunities and Risks of Foundation Models. Stanford Center for Research on Foundation Models (CRFM). arXiv:2108.07258. 2021. arxiv.org/abs/2108.07258
  2. Moor M, Banerjee O, Abad ZSH, et al. Foundation models for generalist medical artificial intelligence. Nature. 2023;616:259–265. doi.org/10.1038/s41586-023-05881-4
  3. Zhou Y, Chia MA, Wagner SK, et al. A foundation model for generalizable disease detection from retinal images. Nature. 2023;622:156–163. doi.org/10.1038/s41586-023-06555-x
  4. Chen RJ, Ding T, Lu MY, et al. Towards a general-purpose foundation model for computational pathology. Nature Medicine. 2024;30:850–862. doi.org/10.1038/s41591-024-02857-3
  5. Regulation (EU) 2024/1689 of the European Parliament and of the Council (Artificial Intelligence Act), Article 3(63). Official Journal of the European Union. 2024. eur-lex.europa.eu/eli/reg/2024/1689/oj
  6. World Health Organization. Ethics and governance of artificial intelligence for health: guidance on large multi-modal models. Geneva: WHO; 2024. www.who.int/publications/b/70584