Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

The replacement question oversimplifies the state of research

The article by Zhihui Zhang and Jing Wang explicitly raises the question in 2024 of whether AI could replace psychotherapists. It describes a broad field: predictive models, monitoring, virtual interventions, general language models, and systems with cognitive-behavioral therapy content. These applications, however, fulfill different tasks and, according to the article, are at different stages of validation. Whoever lumps them together under the replacement question prematurely treats technical language ability, clinical efficacy, and institutional responsibility as one and the same problem.

The editorial position of Dialogatlas is narrower: one can only meaningfully talk about replacement when the specific function, target group, duration, and comparison condition of an application are named. A system can, for example, convey exercises or enable regular interaction. From that it does not yet follow that it can take over diagnostics, long-term support, or professional judgment. The study on XiaoE is relevant precisely for this reason: it does not judge a hypothetical overall AI, but compares a specific offering under controlled conditions with two alternatives.

Three groups, one week, a clearly limited test

He and colleagues recruited 148 young adults with depressive symptoms at a Chinese university. The average age was 18.78 years; 37.2 percent of participants were female. The mean baseline score on the nine-item Patient Health Questionnaire, the PHQ-9, was 10.02; scores ranged from 2 to 19. The study thus examined a group clearly delineated by age, location, and symptom status—not the entire population, nor broadly a group with a diagnosed depressive disorder.

Participants were randomized in a one-to-one-to-one ratio: 49 received access to XiaoE, a dialogue system based on cognitive behavioral therapy, 49 to an e-book, and 50 to Xiaoai, a general conversational assistant. The intervention lasted one week. The study is described as a single-blind, three-arm randomized controlled trial. The primary endpoint was the PHQ-9 after one week and again after one month. In addition, working alliance, acceptance, and usability were included as nonclinical measures.

The result signal is real, but time-limited

In the intention-to-treat analysis, the PHQ-9 scores of the XiaoE group were lower at both measurement time points than in the e-book and general assistant groups. For the one-week time point, the authors report an overall group difference of F(2,136)=17.011, p<0.001, and an effect size of d=0.51. After one month, the difference persisted but was smaller, with F(2,136)=5.477, p=0.005, and d=0.31. In addition, per-protocol analyses, procedures for handling missing data, and sensitivity analyses were performed.

The randomization supports the interpretation that assignment to XiaoE produced a difference compared with the two comparison conditions under these study conditions. The finding does not automatically extend beyond that. The PHQ-9 measures the severity of reported symptoms; lower scores neither demonstrate a lasting change nor equivalence with professionally responsible treatment. The last measurement time point was one month after the start. The authors themselves call for longer interventions, longer follow-up, and comparisons with stronger active control conditions. This is not a side remark but the decisive limit of the study’s scope.

Working alliance without mutual responsibility

The result on working alliance is particularly interesting. The XiaoE group rated it better than the comparison groups; the reported group difference was F(2,145)=3.407 with p=0.04. Acceptance was also higher. For usability, however, no statistically significant difference emerged. XiaoE was therefore not simply superior on all measured product dimensions. Its distinctive feature lay rather in the interaction being perceived as more fitting and more alliance-like.

Yet a questionnaire score on working alliance is not evidence of a human relationship. It shows that participants perceived characteristics such as goal orientation, collaboration, or fit in the interaction, insofar as the instrument used captures these. The system thereby does not assume mutual responsibility and does not develop its own understanding of the person. Above all, the study does not demonstrate that the better alliance caused the lower PHQ-9 score. Both findings occur together; an examination of this mechanism is not reported in the summary. Perceived quality must not be relabeled here as a proven cause of the effect.

The conversational format could work – but through what?

The three study arms allow for a cautious but not unambiguous interpretation. The e-book offered a different mode of delivery, while Xiaoai also enabled a conversation, but without the specific cognitive-behavioral orientation of XiaoE. The fact that XiaoE performed better than both argues against the simple explanation that either mere information or any dialogue suffices. Nevertheless, it remains unclear which component made the difference: the content, the conversation structure, the regularity of use, the fit with the target group, or the interplay of these elements.

Product development in particular tends to mistake the visible surface for the active ingredient. In dialogic software, this surface is the fluent conversation. The XiaoE study suggests a more sober reading: conversational form and professionally structured content can together constitute an effective short-term offering. However, it does not provide a decomposition of the individual components. For research, such a question of mechanism would be more important than the claim that the system already possesses a digital counterpart to human empathy.

Engagement is more than decoration and less than effect

The pilot study by Kien Hoa Ly, Ann-Marie Ly, and Gerhard Andersson extends this point. It examined a fully automated conversational agent for promoting psychological well-being over two weeks and combined a randomized design with qualitative methods. On average, participants opened the application 17.71 times during the study period. The qualitative data also revealed a theme that the authors described as a moderating factor. However, the present summary does not define this feature more precisely, so any further interpretation would be speculative.

The pilot trial primarily supports the assumption that dialogic formats can achieve considerable usage. It does not prove that frequent opening by itself brings about psychological improvements. Rather, the authors viewed the good adherence as a reason for a replication with a larger sample and an active control group. This reflects an important methodological discipline: engagement is a prerequisite for a digital intervention to have any effect at all. But it is neither a clinical endpoint nor a substitute for a robust comparison.

Large language models are a different matter

Zhang and Wang extend the discussion to general language models such as ChatGPT. They refer, among other things, to studies in which such models were able to linguistically recognize and formulate emotional components of hypothetical situations. At the same time, they make clear that this performance is based on pattern recognition and language modeling, not on experienced emotions. A model can therefore produce an empathetic-sounding response without feeling sympathy. Furthermore, a test of linguistic emotion representation is not a clinical efficacy study.

This difference is central to the comparison with XiaoE. The randomized study tests a specifically designed system in a defined intervention. The 2024 contribution, by contrast, discusses a technologically and institutionally much broader future. Its references to short follow-up periods, small study groups, limited long-term memory, algorithmic biases, and necessary human oversight qualify its own far-reaching visions of the future. The efficacy of a specific offering cannot be derived from the capabilities of general models, any more than the suitability of any language-generating system can be derived from XiaoE.

The unanswered questions extend beyond the chat window

The XiaoE results were obtained at a university during the COVID-19 pandemic and involved a very short intervention. This does not automatically impair the internal comparison of the randomized groups, but it limits transferability. The provided sources do not allow us to determine whether comparable results would occur in older adults, in other cultural and institutional settings, with longer-term use, or with more pronounced and complex symptoms. Likewise, the available data do not provide a reliable answer regarding the sustainability of the effect.

This trial also does not address data protection, autonomy, crisis situations, or continuity across many conversations. Zhang and Wang raise such questions, but they do not yet provide an empirical solution. Our editorial conclusion is therefore deliberately narrow: XiaoE provides credible evidence for a short-term difference in young adults with depressive symptoms and for a more positive assessment of the working alliance. It does not provide evidence for a digital relationship partner or for the replacement of professional treatment. Anyone who reads more into the measurement not only goes beyond the data but also confuses convincing conversation design with responsible care.

Sources & further reading