A patient types symptoms into an app or a chatbot and gets back a list of possible conditions and a recommendation: manage at home, see a doctor, or go to the emergency department. This page compares those tools as a safety-first capability matrix: it gathers what peer-reviewed studies actually measured, ties every cell to its source, and names no single best tool — because in a category where a wrong answer can send someone home from a heart attack, a ranking would promise a confidence the evidence does not support. As of July 2026. These tools are a patient-facing form of clinical decision support, and they should be read with the same scrutiny.
How to use this comparison
Two numbers matter for every triage tool, and they behave differently. Diagnostic accuracy asks whether the tool names the right condition; triage accuracy asks whether it gets the urgency right — and whether its errors are safe (sending a low-acuity patient to care they did not need) or unsafe (reassuring a high-acuity patient who needed emergency care). A tool can be mediocre at diagnosis and still be useful at triage, and a tool that looks safe on average can be dangerous in precisely the emergencies that matter. Read every figure against the study behind it, using our guide on how to read an AI validation study.
The pattern that matters most: triage beats diagnosis
Across the literature, one finding recurs: these tools are better at deciding how urgent a case is than at naming the condition. A systematic review of 48 symptom checkers found primary diagnostic accuracy in the range of 19-38%, while triage accuracy ran 48.8-90.1% — and concluded that "reliance upon symptom checkers could pose significant patient safety hazards" 2. The gap is consistent enough to be a design fact of the category, and it points to how the tools should be used: to help decide whether to seek care, rather than to settle what is wrong.
But triage safety carries a hidden trap. A study of 12 tools across 50 vignettes found appropriate triage advice in 57.7% of cases and safe advice overall in 82.6% — yet safety fell as urgency rose, to 71.8% for emergent cases while reaching 87.3% for non-urgent ones 6. Safety dropped exactly where the cost of error is highest. That inversion — cautious with minor complaints, less reliable with emergencies — is the single most important thing to understand before trusting any tool in this category.
Diagnostic accuracy across studies
The numbers below come from different designs — clinical vignettes, real emergency-department data, and a randomized head-to-head — so they answer slightly different questions and should be read side by side rather than pooled.
| Study | Design | Tools | Diagnostic accuracy |
|---|---|---|---|
| Vignette comparison 1 | 8 apps vs GPs | Ada, Babylon, Buoy, K Health, Mediktor, Symptomate, WebMD, Your.MD | GPs 82.1% top-1; best app (Ada) 70.5%; apps span 23.5-70.5% |
| Systematic review 2 | 10 studies, 48 checkers | Multiple | Primary 19-38%; top-3 33-58% |
| ED clinical data 4 | 40 real ED patients | Ada, WebMD, ChatGPT | Top-1: Ada 30%, WebMD 40%, GPT-4 33%, physicians 47%; top-3: Ada 63%, physicians 69% |
| Randomized ED study 5 | 437 patients, double-blind | Ada vs Symptoma | Top-5 correct-or-plausible: Ada 75%, Symptoma 64% |
| Self-triage review 3 | Apps, LLMs, laypeople | Multiple | Apps 25.9-88.0%; LLMs 57.8-76.0%; laypeople 47.3-62.4% |
Two readings deserve emphasis. First, the spread within a single study is enormous — the vignette work found apps ranging from 23.5% to 70.5% on the same tasks 1 — so "symptom checkers" is far too broad a category to carry one accuracy number. Second, even the strongest tools trail clinicians: in the randomized study, Ada clearly outperformed Symptoma, yet "both showed concerning rates of missing potentially life-threatening diagnoses" 5. The general-purpose models are no shortcut here — the self-triage review found large language models weakest of all on self-care cases, at 10.8% accuracy 3, a reminder that fluency is not judgment and that the hallucination risk applies to patient-facing use too.
The self-triage review is worth reading in full for one uncomfortable comparison: it set symptom apps (25.9-88.0%), general models (57.8-76.0%), and ordinary laypeople (47.3-62.4%) side by side 3. The tools beat unaided patients on average, but not by the margin their marketing implies, and even at the urgent end no group cleared roughly three-quarters accuracy on emergency cases 3. "Better than a worried person with a search engine" is a low bar, and one these tools only sometimes clear — which is the honest frame for what a patient-facing triage chatbot is for. It can help a person decide whether to seek care and organize what to say when they do; it has not been shown to replace the judgment of the clinician they reach.
Triage safety across tools
Where diagnosis is measured by whether the condition is named, triage is measured by whether the urgency advice is safe. The vignette study reported the safety of each app's urgency advice against a GP benchmark of 97.0%:
| Tool | Safe urgency advice 1 |
|---|---|
| GPs (benchmark) | 97.0% |
| Symptomate | 97.8% |
| Ada | 97.0% |
| Babylon | 95.1% |
| Your.MD | 92.6% |
| Mediktor | 87.3% |
| K Health | 81.3% |
| Buoy | 80.0% |
The top of that list matched GPs; the bottom sat well below, a spread of nearly 18 percentage points among tools a patient cannot easily tell apart. Real-world data sharpens the point: using actual ED patient records, Ada agreed with physician triage in 62% of cases, but gave unsafe (under-triage) advice in 14% and over-triaged in 24% 4. Over-triage wastes resources; under-triage can kill — and the two are counted separately here for that reason.
Over-triage has a cost too
Unsafe advice gets the headlines, but over-caution exacts its own toll. The 12-tool study found that 51.0% of cases "required additional resource utilisation than that recommended" 6 — a tool that reflexively says "see a doctor" scores well on a safety audit while quietly pushing low-acuity patients into clinics and emergency departments that did not need them. For self-care conditions the effect was starkest, with the great majority of over-cautious cases suggesting an unnecessary in-person visit 6. A tool that is safe only because it escalates everything has answered the safety question by recreating the workload problem the patient opened the app to avoid. This is why the two error types belong on separate lines: the right target is advice that is safe and proportionate, and averaging the two hides whether a tool has reached it.
Vignettes versus real patients
Most of the strongest evidence here rests on clinical vignettes — scripted patient stories fed to each tool under controlled conditions. Vignettes make fair comparison possible, but they can flatter a tool, because a written vignette presents cleaner, better-organized symptoms than a frightened person types at midnight. The studies that used real inputs are accordingly less forgiving: the randomized study drew on 437 actual emergency-department patients 5, and the clinical-data analysis on the messy self-reports of people who had just arrived for care 4. Where a tool's vignette performance and its real-world performance diverge, trust the real-world number, and treat a vignette result as an upper bound rather than a promise.
The regulatory frame
Regulatory status in this category is genuinely mixed. Some symptom assessment apps are marketed as regulated medical devices in Europe; many general chatbots are not positioned as devices at all. Whether a given tool falls under device rules turns partly on its stated purpose and on whether its advice can be independently reviewed rather than simply followed — the principle at the center of the FDA's decision-support guidance 7. For a patient at home, "independent review" is thin — there is no clinician in the loop — which is precisely why the safety-with-urgency inversion above is so consequential. Confirm each tool's current authorization and intended-use claims rather than assuming the category answer.
How to choose for your setting
Whether you are a patient, a clinician recommending a tool, or a system deploying one, ask questions rather than seeking a top pick.
- Diagnosis or triage? These tools are stronger at urgency than at naming the condition 2; match the tool to the job.
- How does it fail on emergencies specifically? The safety-with-urgency inversion 6 means average safety hides the risk that matters — ask for emergent-case performance, not the headline.
- Is there independent, real-world evidence for this exact tool? Vignette results and real-ED results can diverge 45; prefer tools with published data on patients like yours.
- What is its regulatory status and intended use? A tool cleared as a device for a stated purpose carries different assurances than an unregulated chatbot 7.
- Does it advise a safe default? A low threshold to escalate to in-person care is a feature, given the documented under-triage rates.
- Is over-caution acceptable in your context? A tool that escalates freely may be right for a worried patient at home and wrong for a stretched system, because the over-triage it produces 6 lands as avoidable visits somewhere downstream.
How to read this comparison
Four cautions travel with the matrix. First, much of the strongest evidence uses clinical vignettes rather than real patients, and vignettes can flatter a tool by presenting cleaner symptom stories than real people give — where real-ED data exists 45, it tends to be less forgiving. Second, tools update continually, so a result for one version may not hold for the next. Third, "symptom checkers" spans an enormous accuracy range, so no single number represents the category. Fourth, every figure here is dated July 2026, and each tool's evidence and authorization change — reconfirm before relying on any single cell, because a name that led one year can lag the next as tools and their underlying models are rebuilt. The running tally of trial results across clinical AI lives in our clinical AI trial results tracker.
Sources and method
This comparison synthesizes five peer-reviewed evaluations — a vignette comparison against GPs 1, a systematic review of 48 checkers 2, a self-triage review spanning apps, LLMs, and laypeople 3, a real-ED clinical-data study 4, and a randomized double-blinded head-to-head 5 — together with a 12-tool vignette study of triage safety 6 and the FDA's clinical decision support guidance for the regulatory frame 7. Every filled matrix cell is tied to one of these. We revisit this page on a 180-day cycle and whenever a new peer-reviewed study or a change in a tool's authorization lands. For the tools clinicians use to answer their own questions, see our clinical reference AI comparison.