Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
The guiding question lies in the word success
Whether a mental-health conversational system is successful depends first on what it is supposed to achieve. Is its comprehensibility being assessed, its acceptance, the support of adherence, or a change in symptoms? A positive result on one of these levels cannot automatically be transferred to the others. Perceived quality can facilitate use, but it is not yet evidence of a health effect. Conversely, a narrowly limited effect on a symptom scale is not equivalent to a comprehensively helpful conversation.
The three publications at hand therefore address related but not identical subjects. In 2019, Vaidyam and colleagues examined the early research on conversational systems for screening, diagnostics, and treatment of mental disorders. In 2025, Feng and colleagues reviewed controlled studies on specific mental-health outcome measures in 12- to 25-year-olds. In 2020, Le Glaz and colleagues took a broader view of how machine learning and language processing analyze mental-health data. Taken together, they do not provide a comprehensive accounting of AI, but they do offer a useful ordering of different concepts of success.
Ten studies gave rise to great hope in 2019
The lead publication by Vaidyam and colleagues is based on a systematic search from June 2018. PubMed, Embase, PsycINFO, Cochrane, Web of Science, and IEEE Xplore were searched. Studies with a conversational system in a mental-health application context were considered, particularly in people with depression, anxiety disorders, schizophrenia, bipolar disorder, or substance use disorders, as well as in similarly at-risk groups. Of 1,466 records found, eight met the inclusion criteria; two more were added via reference lists. The review was thus based on ten studies.
This small number does not devalue the work. Rather, it precisely describes how young the research field was at the time. Across various studies, the authors reported a high potential for conversational systems. Psychoeducation and support for adherence were particularly highlighted. Satisfaction ratings also consistently came out high. This was a constructive early finding: people could apparently experience such systems as pleasant and usable, and the applications studied were not merely technical demonstrations without any discernible connection to mental health care.
The strongest early finding was acceptability
The positive reception in particular deserves attention. A system that conveys information accessibly or supports people with agreed steps can be practically useful without having to take on an entire professional role. The 2019 review suggests that the conversational form was suitable for such limited tasks. This is more than a side issue, because a technically possible application has little value if it is incomprehensible or consistently avoided.
The conclusion at the time, however, linked high satisfaction with the assumption that the systems could be effective and pleasant tools in psychiatric treatment. From an editorial perspective, this connection must be broken: the ratings primarily support the statement that applications were experienced positively. For efficacy, appropriate comparison groups, defined endpoints, and sufficiently uniform measurements are needed. It was precisely here that the review itself named its limitation. The included studies were heterogeneous, and the authors required standardized outcome reports in order to assess effectiveness more thoroughly.
The more recent meta-analysis refines the promise
Six years later, Feng and colleagues were able to ask a much more specific question: Do AI-supported conversational systems improve measurable mental health outcomes in adolescents and young adults? Their systematic search covered five databases and extended to August 6, 2024. Randomized controlled trials with 12- to 25-year-olds were included. Two people extracted the data independently, a third checked them; the risk of bias was assessed using the Cochrane instrument. In total, 14 articles with 15 studies and 1,974 participants were included in the meta-analysis.
For depressive symptoms, after adjusting for publication bias, a moderate to large pooled effect emerged compared with control conditions: Hedges g was 0.61 with a 95% confidence interval from 0.35 to 0.86. This is a more robust signal of benefit than the earlier satisfaction values, because randomized comparisons and a defined outcome were brought together here. It justifies interest in such systems for depressive symptoms in young people. However, it does not justify a blanket statement about mental health, all age groups, or every form of conversational AI.
Depression does not stand for all mental health outcomes
For generalized anxiety symptoms, stress, positive and negative affect, and psychological well-being, the effects adjusted for publication bias were not significant. This does not mean that an effect is ruled out for all time. It means, more precisely and more importantly: the pooled studies did not demonstrate a statistically significant advantage for these outcome measures. The positive depression finding therefore must not be used as a proxy for a generally effective mental-health conversational system.
In the subgroup analyses, the effect on depressive symptoms was particularly pronounced in subclinical populations; there, Hedges' g was 0.74 with a 95% confidence interval from 0.50 to 0.98. The authors see potential for early interventions in this finding. At the same time, they cite heterogeneous therapeutic orientations of the systems and missing follow-up observations as limitations. On this basis, little can be said about the duration of the effects. The progress compared with 2019 lies not in boundless confirmation, but in a clearer delimitation of the benefit.
Not every language AI even conducts a conversation
The systematic review by Le Glaz and colleagues broadens the perspective to a different technical role. It followed the PRISMA guidelines, was registered with PROSPERO, and searched four databases without time restrictions. Of 327 identified articles, 58 were included and qualitatively analyzed. The studies used machine learning and natural language processing, among other things, to extract symptoms from texts, classify illness severity, obtain indications of psychopathology, or compare the efficacy of treatments. Important data sources were medical records and social media.
Such systems do not necessarily converse with an affected person. They can organize texts in the background, recognize patterns, or prepare information for clinical decisions. The review reports that powerful classifiers were preferred over transparently functioning models. At the same time, the methods frequently confirmed existing clinical hypotheses rather than producing entirely new insights. The authors also point out that social media represent an imprecise population and that language-specific features affect transferability. The common denominator with conversational systems is language as a data form, not automatically the same application or evidence.
The technical role determines the appropriate evidence
These differences have a practical consequence for research and product development. A system for psychoeducation should be measured by whether information is conveyed understandably and correctly and whether the intended target group can handle it. A system that claims to influence depressive symptoms, by contrast, requires controlled comparisons and symptom-related endpoints. An analysis tool for medical records, in turn, demands examinations of its classification performance, its linguistic transferability, and its contribution to professional work. The same umbrella term AI does not make these types of evidence interchangeable.
This ordering is not a mere methodological subtlety. It prevents a product from merging the satisfaction from one study, the symptom reduction from another, and the classification performance of a third technical approach into a single performance promise. None of the three sources supports such an addition. Rather, it is plausible that different systems can be useful in different places: in direct dialogue, in conveying knowledge, in supporting adherence, or in analyzing linguistic data. How well a specific application combines these tasks remains to be examined separately in each case.
Selective confidence is more precise than blanket trust
The development from 2019 to 2025 is neither a story of disappointed euphoria nor a straightforward triumph. The early optimism regarding satisfaction, psychoeducation, and adherence is not refuted by the more recent meta-analysis. It is complemented by a dimension in which a benefit that has been examined in controlled studies emerges for depressive symptoms in young people. Equally clear is that several other outcome measures showed no significant effect and long-term follow-ups are lacking.
Our editorial position is therefore: Mental-health conversational AI should not be judged by its human-like surface, but by the specific task and the corresponding outcome. This is not a defensive stance. On the contrary, it allows us to clearly acknowledge a proven benefit without making it implausible through exaggerated claims. The most interesting perspective currently lies not in the promise of a universal counterpart, but in clearly limited systems whose pleasant use, informative quality, and health impact are each treated as independent findings.
Sources & further reading
- Aditya Vaidyam, Hannah Wisniewski, John Halamka (2019): Chatbots and Conversational Agents in Mental Health: A Review of the Psychiatric Landscape
- Yi Feng, Yaming Hang, Wenzhi Wu (2025): Effectiveness of AI-Driven Conversational Agents in Improving Mental Health Among Young People: Systematic Review and Meta-Analysis
- Aziliz Le Glaz, Yannis Haralambous, Deok-Hee Kim-Dufor et al. (2020): Machine Learning and Natural Language Processing in Mental Health: Systematic Review