Ask a health system whether to deploy an open-weight model it hosts itself or a closed model it reaches through a vendor's API, and the honest answer starts with a refusal to answer in the abstract. The decision is a set of trade-offs, each landing differently depending on what the model will do, whose data it will touch, and who will be accountable when it drifts. This page lays those trade-offs out attribute by attribute, ties each to a primary source, and declines to crown a side. As of July 2026.
Performance is no longer the axis that decides
For two years the reflexive assumption was that closed frontier models were simply better, and that open weights meant a quality compromise. The measured record has closed most of that gap. A peer-reviewed evaluation reported an open-weight model performing on a par with proprietary models on clinical decision-making tasks 2, and on the MedHELM leaderboard the reasoning-tuned open-weight DeepSeek R1 posted a 66% win-rate, leading a field of frontier models as of July 2026 7. A peer-reviewed review of clinical and biomedical information extraction reaches the same practical conclusion from the applied side: the choice is a navigation of trade-offs, with open models adaptable through continued pretraining and fine-tuning and proprietary services carrying usage-based operating cost that can shift as vendors change their terms 1.
The takeaway is liberating and slightly disorienting: because performance no longer separates the options cleanly, the decision moves to axes that procurement, security, and clinical governance own — where the data goes, who validates, and what you can audit.
The trade-off matrix
Read this as a map of tensions, each cell sourced, none scored. Neither column is the "right" one; the right one is the one whose concessions your setting can absorb. As of July 2026 — figures and terms change; reconfirm before you commit.
| Dimension | Open-weight, self-hosted | Closed, proprietary API |
|---|---|---|
| Model weights | Downloadable and inspectable; can be fine-tuned locally 1 | Not released; used as a service 1 |
| Data residency (PHI) | Can run entirely on your infrastructure; data need not leave it 1 | Data is sent to the vendor endpoint under a business associate agreement 3 |
| Measured clinical performance | On a par with proprietary in a peer-reviewed evaluation; open reasoning models lead MedHELM as of July 2026 27 | Historically strong; parity now contested, not assured 2 |
| Validation & change-control burden | You own validation, updates, and the change-control plan for the version you run 5 | Vendor updates the model outside your control; silent version changes can shift behavior 1 |
| Transparency / auditability | Weights and (sometimes) training details inspectable, but data and copyright often opaque 4 | Limited disclosure; the 2024 index found systemic opacity across developers 4 |
| Operating cost shape | Infrastructure plus engineering effort you carry | Usage-based cost set by the vendor and subject to change 1 |
| Governance / accountability | Rests with the deploying hospital 3 | Rests with the deploying hospital 3 |
Two rows deserve unpacking because they are the ones most often flattened in a sales conversation.
Data residency is a capability, never a guarantee. A self-hosted open-weight model can keep protected health information inside your walls, which is why de-identification and data-residency questions so often push teams toward open weights. But open weights do not by themselves make a deployment private or compliant — a poorly secured self-hosted model can leak as surely as a misconfigured API, and a closed API under a proper business associate agreement can be a defensible choice. The weights set what is possible; your controls determine what is true.
The validation burden shifts rather than disappears. With a self-hosted model, you own every update, which means you also own the predetermined change control plan and the local validation of each version. With a closed API, the vendor can change the model beneath you — a convenience that removes maintenance work and introduces model drift you did not schedule and may not be told about. The FDA record shows how immature this discipline still is even among cleared devices: among 2024 machine-learning authorizations, only 29.2% reported both sensitivity and specificity and a change-control plan appeared in only 16.7% of summaries 5. Whichever column you pick, the change-control gap is yours to close.
Transparency is measurable — and thin on both sides
It is tempting to equate "open" with "transparent." The evidence complicates that. The 2024 Foundation Model Transparency Index scored leading developers at 58 out of 100 on average and found "sustained and systemic opacity" across developers on copyright status, data access, data labor, and downstream impact 4. Open-weight releases tend to score better on model-access dimensions, yet the hardest questions for a clinical buyer — what the model was trained on, whether that data was licensed, how it behaves downstream — remain under-disclosed regardless of the license. A foundation model you can download is more inspectable than one you cannot; it is still not fully legible.
There is a regulatory backdrop worth stating plainly. As of the most recent peer-reviewed taxonomy, none of the FDA-authorized medical devices use large language models at all 6 — the cleared field is quantitative image analysis. So a hospital deploying a clinical LLM, open or closed, is generally operating outside the cleared-device pathway, under its own governance rather than a clearance. That raises the stakes on internal validation for both columns.
Governance does not travel with the weights
The most consequential row in the matrix is the one where the two columns read the same. WHO's guidance on large multi-modal models assigns responsibilities across the model lifecycle to developers, providers, and deployers, and it does not let the choice of an open or closed model relocate accountability 3. Whoever puts the model in front of a clinician owns the duty to validate it locally, monitor it, and act when it fails. Open weights can make that duty easier to discharge — you can inspect and test what you run — but they do not discharge it for you, and a closed API does not hand the duty to the vendor. The hospital is accountable either way.
How to choose for your setting
Questions, never a recommendation. The answers are yours, and they will differ by service line and by country.
- Where must the data physically stay? If regulation or policy requires PHI to remain on your infrastructure, that constraint points one way before any performance comparison 13.
- Can you carry the validation and change-control burden? Self-hosting asks for engineering and monitoring capacity; an API trades that for dependence on a vendor's release cadence 5.
- How much must you be able to audit? If you need to inspect and freeze the exact version behind a clinical decision, open weights make that possible; a changing API does not 4.
- Is measured performance actually the constraint? Given the parity evidence, confirm whether performance genuinely separates your candidates or whether the decision is really about residency and control 27.
- Who signs for the outcome? Name the accountable owner before deployment. WHO's framing is blunt: that owner is you, whichever column you chose 3.
For the moving performance picture, our LLM medical benchmark tracker follows the leaderboards these claims rest on; for the layer above the model, see imaging AI marketplaces compared.
Sources and method
This trade-off matrix draws its parity evidence from a peer-reviewed clinical evaluation 2 and the MedHELM leaderboard 7, its applied framing from a peer-reviewed review of clinical information extraction 1, its transparency figures from the 2024 Foundation Model Transparency Index 4, its regulatory reality from two analyses of the FDA record 56 and the FDA's own device list 8, and its governance frame from WHO's guidance on large multi-modal models 3. We present attributes neutrally and name no winner; inclusion of a model or approach here is not an AIMOCS endorsement. Figures such as transparency scores and leaderboard positions are perishable; we revisit this page on a 180-day cycle and whenever a new transparency index, a clinical head-to-head, or revised FDA/WHO guidance lands. As of July 2026.