Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
A broad class of interventions, not a uniform technology
The lead publication by Yuhao He and colleagues examined so-called conversational agent interventions for various psychological problems. The research group aimed to summarize the characteristics of such interventions, examine their efficacy, and identify statistically significant moderators. Included were randomized controlled trials in which a conversational agent was compared with any control condition. Among the outcomes considered were depressive and generalized anxiety symptoms, specific anxieties, stress, general distress, well-being, psychosomatic complaints, and positive and negative affect.
This range alone limits any blanket statement. The common term encompasses interventions for different target groups, problems, and outcome measures. The meta-analysis thus does not evaluate any single application, and certainly not automatically every current generative AI system. From an editorial perspective, this distinction is central: an average effect of a broad research category is neither a seal of quality for a specific product nor evidence that different conversational systems are interchangeable.
32 studies provide a robust but limited overview
He and colleagues searched six scientific databases from their inception to October 30, 2021, and updated the search until May 1, 2022. Of 6900 identified records, 32 studies with a total of 6089 participants were included. Two people extracted the data independently, and a third checked them. The work followed the PRISMA guidelines and pooled the results using both random-effects and fixed-effects models. Hedges' g served as the effect measure.
The study design is essential to the guiding question: only randomized controlled comparisons were considered. In this way, the work goes beyond mere satisfaction scores or voluntary usage statistics. Nevertheless, a meta-analysis can only answer the questions that its primary studies actually investigated. It can pool participant counts, but it cannot retroactively generate missing follow-ups, different interventions, or uncollected care outcomes.
The clearest common finding lies in the short term
Compared with the respective control conditions, the conversational interventions showed statistically significant short-term effects on all distress measures reported by the meta-analysis. For depressive and generalized anxiety symptoms, Hedges g was 0.29 in each case. Effects were also reported for specific anxiety symptoms of 0.47, for quality of life or well-being of 0.27, for general distress of 0.33, and for stress of 0.24. For symptoms of mental disorders, g was 0.36, for psychosomatic complaints 0.62, and for negative affect 0.28.
These results argue against the claim that digital conversations are fundamentally ineffective. However, they demonstrate differences in standardized measurements, not automatically a change that is noticeable in everyday life or clinically meaningful. The meta-analysis reports statistical significance; a separate threshold for clinical relevance does not emerge from the provided source evidence. This is precisely where correct statistics are often turned into an overblown narrative: a demonstrable group difference is treated as proof of comprehensive suitability for care.
After the intervention, the statement becomes vague
For most outcomes considered over the longer term, the effects were not statistically significant. The reported long-term values ranged from g = -0.04 to 0.39. This does not prove that no effect persists. But it means that this meta-analysis could not establish a statistically confirmed long-term benefit for most mental health outcome measures.
This difference is more important for product development and care than a blanket debate about “efficacy.” A system can show measurable benefit during or immediately after an intervention without that benefit remaining stable later. Our assessment is therefore: the clearest evidence at present is short-term. Anyone who promises lasting stabilization, prevention of later deterioration, or sustainable relief goes beyond the evidence reported here.
Empathetic responses are a moderator, not proof of a relationship
The lead publication identifies personalization and empathetic responses as important facilitating factors. In addition, longer interaction duration was associated with larger pooled effect sizes. These are relevant indications for the design of digital conversational offerings: not only the provision of a dialog window seems decisive, but also how responsive and attentive the interaction is experienced or implemented.
However, a moderator analysis is not a causal experiment on individual product features. The fact that longer interactions are associated with larger effects does not prove that merely extending a conversation produces the effect. Likewise, an empathetically designed response does not demonstrate a human relationship or a professional working alliance. Rather, it is plausible that design, usage intensity, and other characteristics vary together. For developers, the findings are therefore hypotheses with empirical weight, but not a ready-made construction manual.
The older evidence was more cautious – for good reason
The systematic review by Alaa Abd-Alrazaq and colleagues from 2020 found twelve studies on eight outcome measures. It rated the evidence for improvements in depression, distress, stress, and acrophobia as weak. For subjective psychological well-being, there was no statistically significant effect; for anxiety as well as positive and negative affect, the results were contradictory. As reasons for the cautious conclusion, the authors cited, among other things, few studies per outcome measure, a high risk of bias, contradictory results, and a lack of evidence for clinically meaningful effects.
The evidence base on safety was also narrow. Only two included studies examined it; no adverse events or harms were reported there. However, the absence of reports in two studies is not comprehensive evidence of safety. The comparison with the 2023 meta-analysis shows progress in scope and randomized evidence, but does not eliminate every earlier uncertainty. In particular, the questions of clinical significance and longer-term course remain clearly evident.
In young people, the benefit is concentrated on depression
The meta-analysis published in 2025 by Yi Feng and colleagues narrows the focus to 12- to 25-year-olds and explicitly AI-driven conversational agents. It includes 14 articles with 15 randomized studies and a total of 1974 participants. After correction for publication bias, a moderate-to-large effect of g = 0.61 was shown for depressive symptoms. In subclinical populations, the corresponding value was 0.74.
For generalized anxiety symptoms, stress, positive and negative affect, and psychological well-being, the effects adjusted for publication bias were not statistically significant. The authors cite different therapeutic orientations of the systems and missing follow-up observations as key limitations. Thus, the more recent analysis extends the lead publication but contradicts a blanket application of its results: In young people, the benefit does not appear as a uniform effect across different areas of distress, but primarily as a finding on depressive symptoms.
Evidence is lacking that links symptom value to relief of the care system
He and colleagues conclude from their results that broader implementation is warranted and link the technology to the prospect of conserving human resources and distributing care services better. This conclusion is more far-reaching than the results reported in the source findings. The meta-analysis primarily aggregates mental health outcome measures. Saved working time, changed utilization of professional services, or better resource distribution are not quantified in the reported results. Nor do easy accessibility or acceptance substitute for such evidence.
This is the core conflict: A short-term symptom difference can be valuable for users, but it does not answer which task a system reliably performs in a real care chain. It remains open whether digital conversations actually reduce human work, merely generate additional demand, or shift tasks. The three meta-analyses do not resolve this organizational question.
Our editorial position is therefore deliberately narrower than the most optimistic conclusion of the lead publication. Digital conversational agents should be assessed according to what has been tested for a specific target group, target outcome, and time span. The existing research evidence allows limited statements about short-term effects, especially for depressive symptoms. It justifies neither a general long-term promise nor the assumption that statistical efficacy automatically brings with it safety, clinical relevance, and relief of the care system.
Sources & further reading
- Yuhao He, Li Yang, Chunlian Qian (2023): Conversational Agent Interventions for Mental Health Problems: Systematic Review and Meta-analysis of Randomized Controlled Trials
- Yi Feng, Yaming Hang, Wenzhi Wu (2025): Effectiveness of AI-Driven Conversational Agents in Improving Mental Health Among Young People: Systematic Review and Meta-Analysis
- Alaa Abd‐Alrazaq, Asma Rababeh, Mohannad Alajlani (2020): Effectiveness and Safety of Using Chatbots to Improve Mental Health: Systematic Review and Meta-Analysis