The strongest argument for mental health chatbots is not that they work well. It is that in most of the world, the waiting list is measured in months and the alternative on offer is nothing at all.
That framing is honest, and it is also incomplete — because "nothing" is not the only alternative, and a poorly designed tool can leave someone worse off than no tool would have. The comparison that matters is not chatbot versus therapist. It is chatbot versus whatever that specific person would otherwise do, which might be a friend, a helpline, a search engine, or silence.
Holding both of those thoughts at once is the only useful way to think about AI in mental health, and most writing on the subject manages only one.
The Distinction That Determines Everything
There are two fundamentally different products being discussed under the same heading, and conflating them is the source of most confusion.
Structured therapeutic tools deliver a defined intervention — usually derived from cognitive behavioural therapy — through a constrained conversational interface. They follow a protocol. What they can say is bounded. Several have been evaluated in randomised trials, and some have received regulatory designations acknowledging their potential.
Open-ended conversational companions are general-purpose systems that will discuss anything, in any direction, for as long as the user wants. They were not designed as clinical tools, were not evaluated as clinical tools, and have no protocol constraining what they say.
The evidence base attaches almost entirely to the first category. The public conversation, the media coverage, and the documented harms attach almost entirely to the second.
What the Evidence Actually Shows
For structured, protocol-based tools, published trials have generally found modest short-term improvements in symptoms of depression and anxiety compared with waitlist or minimal-support control groups.
Several caveats come with that, and they are substantial:
- Effect sizes are modest, and typically smaller than those found for therapy delivered by a person.
- Comparison groups are often weak. Beating a waiting list is a lower bar than beating an active treatment.
- Trial populations are unrepresentative — frequently university students, frequently people with mild to moderate symptoms, frequently excluding anyone at meaningful risk.
- Follow-up is short. Whether benefits persist beyond a few weeks is largely unstudied.
- Attrition is high. Many participants stop using these tools quickly, and dropouts are not always fully accounted for.
- Much of the research is funded by the developers, which is common in this field and worth knowing.
The reasonable summary is that structured digital tools appear to help some people with mild to moderate difficulties, somewhat, in the short term. That is a real finding. It is not what the marketing claims.
What Happened When a Chatbot Replaced a Helpline
In 2023, the National Eating Disorders Association in the United States moved to replace its human-staffed helpline with a chatbot. Within a short period, reports emerged that the system was providing weight-loss and calorie-restriction guidance to people contacting it about eating disorders — advice that is actively dangerous for that population. The chatbot was withdrawn.
The failure is worth understanding precisely, because it was not a technical malfunction. The system produced responses that would be unremarkable in a general wellness context and were harmful in the specific context of the people it was serving.
The lessons generalise:
- Context determines whether advice is safe. Generic guidance delivered to a vulnerable population is not neutral.
- Replacing a human service is different from supplementing one. The people who would have reached a trained volunteer instead reached a system that could not recognise what it was dealing with.
- Testing must involve the actual population, including people in acute distress, not general users.
- Someone must be accountable for what the system says, with the authority to shut it down quickly.
Companion Chatbots and Young People
The most serious safety concerns in this field currently involve general-purpose companion applications and adolescent users.
Litigation in the United States has raised allegations concerning the role of companion chatbots in cases involving young people who came to serious harm. Professional bodies in psychology have issued advisories about adolescent use of these systems, and several jurisdictions have moved toward legislation restricting the provision of therapy-like services by artificial intelligence without clinical oversight.
The specific properties that make these systems risky for this group:
- They are designed to be engaging, and engagement optimisation and clinical benefit are not the same objective.
- They do not reliably recognise deterioration, because they were not built to.
- They are available continuously, which can substitute for rather than support human connection.
- They tend toward agreement, and a system that validates whatever a distressed person says is not providing what a therapist would.
- No one is accountable in the way a clinician is accountable.
Anyone with responsibility for a young person using these applications should know what they are — general conversational software, not a clinical service — and should treat prolonged reliance on one for emotional support as something worth a conversation rather than something to ignore.
Early Detection: Promising, Unproven
Research into detecting mental health changes from passively collected data — phone usage patterns, movement, sleep timing, typing behaviour, speech characteristics — is genuinely interesting and genuinely immature.
The premise is reasonable. Depression and mania affect sleep, movement, and social contact, and phones record all three continuously. Studies have found associations between these signals and mood states.
What has not been established:
- That these signals predict deterioration reliably enough at the level of an individual to act on.
- That intervening on such a prediction improves outcomes.
- That models transfer across populations, cultures, and devices.
- That the surveillance required is acceptable to the people being monitored.
Suicide risk prediction models built from health records deserve a specific mention. Several have been developed and evaluated, and their performance in identifying individuals at risk has generally been modest. Given the gravity of both error types — missing someone at risk, and flagging someone who is not — the ethical bar for deployment is extremely high, and current performance does not clearly clear it.
What These Tools Cannot Do
Being specific about the boundary is the most useful thing this article can offer.
- Assess risk reliably. Judging whether someone is safe involves things a text interface cannot perceive.
- Repair a rupture. A significant part of what makes therapy effective is the experience of a relationship surviving conflict and misunderstanding. There is no relationship to repair here.
- Hold a person in mind. A therapist thinks about a patient between sessions. That mattering to someone is itself therapeutic.
- Be accountable. A clinician is licensed, supervised, and answerable. An application's terms of service disclaim responsibility explicitly.
- Recognise what is not said. Much of clinical assessment is about hesitation, avoidance, and the thing a person circles without naming.
- Escalate meaningfully. Displaying a helpline number is not the same as a clinician arranging care.
Where They Genuinely Help
None of this means these tools are worthless. Used as a supplement rather than a substitute, several applications are defensible:
- Practising skills between therapy sessions — thought records, exposure exercises, behavioural activation.
- Psychoeducation, explaining what a condition is and what treatment involves.
- Structured journaling with prompts, for people who find open journaling difficult.
- Something at three in the morning, when no service is open and the alternative is lying awake alone.
- A first step for someone not yet ready to talk to a person, which may lead to seeking care.
- Administrative relief for clinicians, freeing time for patients rather than notes.
That last one may be the most valuable contribution in the whole field, and it involves no patient-facing chatbot at all.
Questions Worth Asking About Any App
- Is it a regulated medical device, or a wellness product? The difference is substantial and rarely advertised.
- What evidence exists, and who funded it?
- What does it do when someone is in crisis? Test this before recommending it to anyone.
- Where does the data go? Mental health data is among the most sensitive information a person has, and consumer apps often sit outside health privacy law.
- Can it be deleted? Completely, on request.
- Is it a supplement or a replacement, and does the marketing make that clear?
- Who is accountable if it causes harm?
The Position Worth Holding
Access to mental health care is genuinely inadequate almost everywhere, and that gap is real rather than rhetorical. Tools that reach people who would otherwise receive nothing have a legitimate place.
But the deficit is a reason to build carefully, not a reason to accept whatever is offered. "Better than nothing" is a claim that has to be demonstrated for a specific tool and a specific population — and the eating disorder helpline case shows that it is not automatically true.
The most defensible use of AI in mental health today is as scaffolding around human care: helping people practise between sessions, reducing the administrative load on the clinicians who are already scarce, and lowering the barrier to a first conversation with a person.
The least defensible is using it as a cheaper replacement for the human care that is missing. That is not expanding access. It is offering something different and calling it the same thing.

0 Comments