No use of AI in healthcare reaches more people, with less oversight, than the mental-health chatbot. Tens of millions of people have typed a private worry into a conversational app, and the market presents a smooth surface: an always-available, low-cost listener. Underneath, three very different products are being blurred into one, and telling them apart is the whole task. This guide separates them, ties each to its primary evidence, and treats safety as the first question rather than the last. As of July 2026.
A note before the evidence: a chatbot is no substitute for professional care, and none of these tools is a crisis service. If you or someone else is in crisis or at risk of harm, contact local emergency services or a recognized crisis line.
Three categories the market blurs
The single most useful thing you can do is refuse to let these be one thing.
- Structured, purpose-built tools. Chatbots built to deliver a specific protocol — usually cognitive behavioral therapy — with scripted guardrails and, in the strongest cases, trial evidence. Woebot and Therabot sit here.
- General-purpose LLMs acting as therapists. Consumer chatbots and companion apps built on a general foundation model, not designed for care, that users nonetheless treat as a counselor. This is where the documented harms concentrate.
- Wellness apps. Mood trackers, journaling, and meditation tools that market themselves as "wellness" precisely to stay clear of treatment claims and the regulation that follows them.
The trial evidence below belongs almost entirely to the first category. Most of the safety failures belong to the second. Reading a result from one as if it applied to another is the error the whole field turns on.
The appeal is genuine, and worth stating plainly rather than dismissing. The world has a severe shortage of mental-health workers — the shortage that motivated chatbots in the first place 3 — and a tool that is available at 3 a.m., costs little, and carries no waiting list or perceived stigma reaches people who would otherwise reach no one. That is a real good, and it is exactly why the safety questions matter so much: the same properties that make these tools accessible, always-on and unsupervised, also remove the human who would notice when a conversation is going wrong.
What the trial evidence actually shows
The table sets out the strongest studies with their designs. Read each result next to the design that produced it. As of July 2026.
| Study | Design | What it measured | Headline result |
|---|---|---|---|
| Woebot RCT 1 | 70 young adults, 2 weeks, vs information-only control | Depression (PHQ-9), anxiety | Significant drop in depression vs control |
| Therabot RCT 2 | 210 adults, 4 weeks, generative AI vs waitlist, clinician-monitored | Depression, anxiety, eating-disorder concerns | ~51% / 31% / 19% symptom reduction |
| Chatbot meta-analysis 3 | 12 studies pooled | Multiple mental-health outcomes | Weak evidence; safety rarely assessed |
The purpose-built evidence is genuine and worth taking seriously. An early randomized trial of the structured CBT chatbot Woebot found that 70 young adults who used it for two weeks significantly reduced their depression symptoms on the PHQ-9 compared with an information-only control group 1. More recently, the first randomized trial of a generative-AI therapy chatbot, Therabot, reported that across 210 adults, four weeks of use produced meaningful symptom reductions — on the order of 51% for depression, 31% for anxiety, and 19% for eating-disorder concerns relative to a waitlist 2. That is a landmark, and it points to real potential.
It also carries a condition that is easy to skip past: the Therabot trial ran under clinician monitoring, with human oversight available throughout, and its authors are explicit that supervision remains necessary 2. The human-in-the-loop was part of the intervention, which means the result speaks to a supervised tool rather than an autonomous one.
And when the whole literature is pooled, the picture cools. A meta-analysis of 12 studies found only weak evidence that chatbots improve depression, distress, stress, and acrophobia, no significant effect on subjective psychological wellbeing, and conflicting results for anxiety — with formal safety assessed in only a small minority of the studies 3. The honest summary is a promising but thin evidence base, measured mostly over short windows, in which safety has been the least-studied outcome.
The safety failures that matter
The gap between a supervised, purpose-built tool and a general-purpose chatbot acting as a therapist is where the danger lives, and it has now been measured. Independent testing of leading language models and commercial therapy bots found they answered only about half of clinical prompts appropriately — with one widely used therapy bot appropriate just 40% of the time — and that they expressed stigma toward conditions such as schizophrenia and alcohol dependence more than toward depression 4. Most seriously, the models mishandled expressed suicidal ideation: prompted by a user who paired losing their job with a question about tall bridges, a therapy bot responded with a list of specific bridges rather than recognizing the risk 4. The authors' conclusion is in their title — expressing stigma and inappropriate responses prevents these systems from safely replacing mental-health providers 4.
Two properties of general-purpose models drive these failures. The first is sycophancy: a model tuned to be agreeable can validate a user's distorted or dangerous thinking rather than challenge it — the opposite of what a therapist does. The second is hallucination — confident, fluent output that has no grounding in fact or clinical judgment. A clinical LLM deployed for care needs guardrails, escalation paths, crisis-detection routing, and evaluation against exactly these failure modes; a consumer chatbot repurposed as a counselor has none of them by default.
The concern is sharpest for young people, who adopt these tools fastest and are most exposed to their failures. The American Psychological Association's 2025 health advisory grew directly out of this: it followed the APA meeting the Federal Trade Commission over chatbots impersonating mental-health professionals, and it urges clear, repeated disclaimers that a chatbot's output does not substitute for professional care 7. A companion app that a teenager treats as a confidant is operating in the highest-risk part of this field with the least oversight.
The regulatory response is moving fast
Regulators and professional bodies spent 2025 catching up, and the direction is clear. The table tracks the key actions. As of July 2026.
| Actor | Action | Date | Source |
|---|---|---|---|
| US FDA | Digital Health Advisory Committee meets on generative-AI mental-health devices | 6 Nov 2025 | 5 |
| Illinois | WOPR Act bars AI from independently providing therapy | Aug 2025 | 6 |
| APA | Health advisory on generative-AI chatbots and wellness apps | 2025 | 7 |
On 6 November 2025 the FDA's Digital Health Advisory Committee convened specifically on generative-AI-enabled digital mental-health devices, taking a hypothetical prescription language-model therapy chatbot for major depressive disorder as its worked example — a strong signal that the agency intends to treat such tools as regulated devices 5. Crucially, the FDA has not approved any generative-AI mental-health tool as of July 2026. In August 2025, Illinois became the first US state to bar AI from independently providing therapy or psychotherapy through the Wellness and Oversight for Psychological Resources Act, permitting AI only as administrative or supplementary support to a licensed professional, with fines up to $10,000 6. And the American Psychological Association issued a 2025 health advisory urging that AI systems carry clear, prominent disclaimers that their output does not substitute for professional care, and that any system offering health information either ensure its accuracy or warn repeatedly that it may be wrong 7.
Whether a given chatbot is even a regulated device turns in part on the FDA's clinical decision support line: whether a clinician can independently review the basis for its output, or whether a user relies on it directly 8. Many consumer tools sidestep the question by branding themselves as wellness rather than treatment — which is exactly why the wellness category deserves the same scrutiny as the others. The same underlying model can be framed as a "mental-fitness companion" to stay outside the device pathway or as a "treatment" to enter it, and the words in the marketing, rather than anything in the software, often decide which regime applies. For a user in distress the distinction is invisible: the app feels like help regardless of how its maker has classified it, which is why disclosure and honest labeling carry so much of the safety burden here. Our global AI-in-health regulation tracker follows these instruments as they change. Because the legal position varies by jurisdiction and is moving quickly, confirm the current rules with qualified counsel before deploying or recommending any tool.
How to read this field
Five cautions travel with everything above. First, keep the three categories apart: trial evidence for a supervised, structured tool tells you nothing about a general-purpose chatbot a user adopts on their own. Second, supervision is doing work: the strongest positive trial ran under clinician monitoring, so its result is evidence for a supervised tool, and the discipline for reading such claims is in our guide on how to read an AI validation study. Third, safety is the least-measured outcome, so the absence of reported harm in a study is weak reassurance. Fourth, the failure modes are specific and severe — stigma, sycophancy, and mishandled crisis prompts — and they are the first thing any deployment must be evaluated against. Fifth, the rules are changing under your feet: a tool that is permitted today may be restricted tomorrow, and a jurisdiction that allows unsupervised AI therapy now may follow Illinois in barring it. The throughline is that the strongest positive evidence and the clearest safety practice both point the same way — toward AI that supports a clinician rather than stands in for one. For the adjacent setting where these tools most often surface first — the front line of general practice — see our AI in primary care guide.
Sources and method
This guide is built on the mental-health chatbot evidence base: the Woebot randomized trial 1, the Therabot generative-AI trial 2, and the pooled meta-analysis of chatbot effectiveness and safety 3, with the independent safety evaluation of language models as therapists supplying the harm evidence 4. The regulatory frame draws on the FDA Digital Health Advisory Committee meeting on generative-AI mental-health devices 5, the Illinois WOPR Act 6, the APA health advisory 7, and the FDA's clinical decision support guidance 8. Every figure is tied to the primary source cited beside it. We revisit this page on a 180-day cycle and whenever a regulator acts, a state enacts a law, or a new trial reports. Dates and statuses are current as of July 2026. Nothing here is clinical, legal, or crisis advice.