Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

The wrong yardstick

Public debate likes to jump from a fluent answer to a big systemic question: If a machine sounds understanding, can it then replace psychotherapists? The very form of this question shifts the burden of proof. Linguistic plausibility, perceived empathy, and a measured symptom change are three different things. None of these measures automatically stands for the others. Even more so, a statistical advantage over a control condition does not demonstrate equivalence with a trained professional.

This is not a semantic objection but the core of the research problem. A professional treatment includes, among other things, continuity of treatment, situational judgment, and dealing with complex changes. The available meta-analyses, by contrast, examine defined psychological endpoints in time-limited studies. From their results, benefit for specific tasks can be derived. A general proof of replacement does not follow from that. In our editorial assessment, precisely this distinction should determine the discussion.

The meta-analytic core of 2023

Li, Zhang, Lee and their co-authors searched twelve databases for experimental studies on AI-based dialogic agents that had been published before May 26, 2023. Of 7,834 entries found, 35 studies met the criteria for the systematic review. For the meta-analysis, 15 randomized controlled trials were used. This distinction is important: The broader review describes the research field, while the pooled efficacy estimates are based on a considerably smaller randomized core.

The lead publication examines not only whether such systems change psychological target variables. It also asks which characteristics are associated with stronger effects and with the user experience. In doing so, it goes beyond a mere yes/no assessment. At the same time, the systems, target groups, and endpoints studied remain diverse. The meta-analysis condenses this diversity into standardized effect sizes; however, it does not transform the different applications into a uniform product or a uniform intervention.

Two positive findings and a telling gap

For depressive symptoms, the meta-analysis reports a significant reduction with Hedges g of 0.64 and a 95% confidence interval from 0.17 to 1.12. For psychological distress, Hedges g was 0.70; the interval ranged from 0.18 to 1.22. These are positive pooled findings. However, the relatively wide intervals also show that the precise magnitude of the effects remains uncertain. Moreover, statistical significance does not answer the question of long-term stability or clinical significance in individual cases.

In contrast, no significant advantage was found for general psychological well-being. The estimate was Hedges g of 0.32, and the confidence interval, from -0.13 to 0.78, also included zero. This difference is particularly instructive: fewer depressive symptoms or less distress is not the same as comprehensively improved well-being. The evidence thus points to a target-variable-specific effect, not a blanket improvement in mental health.

Larger effects are not yet a blueprint

In the analysis, the effects were stronger when the conversational agents were multimodal or generative in design, were integrated into mobile or instant messaging applications, or targeted clinical, subclinical, and older populations. Such moderator findings are interesting for product development. They suggest that access route, presentation form, technical architecture, and target group are not merely packaging but can be related to the outcomes.

However, they must not be read causally. The meta-analysis identifies associations between study characteristics and effect sizes; it does not prove that, for example, multimodality itself causes the improvement. Other characteristics of the respective studies and applications may underlie these differences. Therefore, a stronger estimate does not imply a reliable development rule. From an editorial perspective, we consider it premature to derive a supposedly optimal AI companion from these subgroups.

A good conversation experience is not yet a mechanism of action

The systematic review describes three factors as particularly important in shaping the user experience: the quality of the human-AI relationship experienced as therapeutic, engagement with the content, and effective communication. This is an important finding for conversation design. A system that comes across as incomprehensible, inappropriate, or lacking in connection is unlikely to be experienced as a helpful counterpart. The study thus shows that conversational quality is not a side issue for users.

However, perceived relationship quality is not yet evidence that the same relationship caused the measured symptom change. Nor is it to be equated with human empathy. A language system can recognize emotional patterns and generate appropriate formulations without experiencing feelings. The lead publication consequently calls for further research into the underlying mechanisms of action. This is precisely where a central knowledge limit lies: we sometimes see changes, but do not yet know sufficiently how they come about.

Among young people, the pattern becomes sharper

The meta-analysis by Feng, Hang, Wu, and colleagues from 2025 narrows the population to 12- to 25-year-olds. Five major databases were searched up to August 6, 2024. 14 articles with 15 randomized studies and a total of 1,974 participants were included. Two people independently extracted the data, a third checked them; the risk of bias was assessed using the Cochrane instrument. However, the provided abstract does not report the result of this quality assessment.

After correction for publication bias, an effect of Hedges g equal to 0.61 with a 95 percent confidence interval from 0.35 to 0.86 was found for depressive symptoms. In subclinical groups, the estimate was 0.74. For generalized anxiety symptoms, stress, positive and negative affect, and psychological well-being, the corrected effects were not significant. Thus, the more recent analysis does not simply confirm a general benefit, but focuses the positive evidence even more clearly on depressive symptoms.

The authors identify different therapeutic orientations of the systems and missing follow-up observations as central limitations. The statement on early support for depression is therefore an indication of potential, not evidence of lasting effect. Also noteworthy is the agreement with the older meta-analysis on well-being: even in the more narrowly defined young population, no significant advantage can be shown for this broad goal.

The replacement question outpaces its own evidence

In 2024, Zhang and Wang explicitly raise the question of whether AI could replace psychotherapists. Their contribution broadens the perspective to scalability, constant availability, possible openness toward machines, and support for overloaded care systems. At the same time, it addresses long-term memory problems, algorithmic biases, data protection, the lack of genuine emotional experience, and the necessity of human oversight. Its conclusion is therefore more cautious than the pointed title: AI is meant to complement human work, not displace its central elements.

Methodologically, it remains crucial that the provided evidence text does not describe a direct randomized equivalence comparison between an AI system and psychotherapists. Even a hypothetical support conversation used there can illustrate a possible form of interaction, but cannot demonstrate efficacy. The claim of an impending replacement thus goes further than the presented evidence. It blends technical capabilities, presumed care advantages, and clinical results into a prognosis that has not yet been tested.

The honest product description begins with the endpoint

For research and product development, this implies an uncomfortably precise language. A system with evidence for reducing depressive symptoms should not be described as generally effective for mental health. A short-term study endpoint says nothing about lasting stability. A positive user experience does not replace a safety check, and a moderator relationship is not a validated design specification. The lead publication explicitly names long-term effects, mechanisms of action, and the safe integration of large language models as open research tasks.

Our editorial position is therefore neither technology-hostile nor euphoric: Conversational AI can have a limited, empirically testable function. It is precisely this limitation that makes serious development possible. The relevant unit is not “the digital therapist,” but a concrete application for a defined target group, target outcome, and duration. As long as studies do not examine direct equivalence with professional treatment, replacement is the wrong category. The robust finding is narrower – and more useful: Some systems change some symptoms, while many more far-reaching efficacy promises remain open.

Sources & further reading