Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
Article history
- First published
- Last substantive revision
The substitution question oversimplifies the state of research
The contribution by Zhihui Zhang and Jing Wang explicitly asks in 2024 whether AI could replace psychotherapists. It describes a broad field: predictive models, monitoring, virtual interventions, general language models, and systems with cognitive-behavioral therapy content. These applications, however, fulfill different tasks and, according to the contribution, are at different stages of validation. Those who lump them together under the substitution question prematurely treat technical language ability, clinical efficacy, and institutional responsibility as one and the same problem.
The editorial position of Dialogatlas is narrower: substitution can only be meaningfully discussed once the specific function, target group, duration, and comparison condition of an application are named. A system might, for instance, convey exercises or enable regular interaction. It does not follow from this that it can take over diagnostics, long-term support, or professional judgment. The study on XiaoE is relevant precisely because it does not judge a hypothetical overall AI but compares a specific offering under controlled conditions with two alternatives.
Three groups, one week, a clearly limited test
He and colleagues recruited 148 young adults with depressive symptoms at a Chinese university. The average age was 18.78 years, and 37.2 percent of participants were female. The mean baseline score on the nine-item Patient Health Questionnaire, the PHQ-9, was 10.02; scores ranged from 2 to 19. The study thus examined a group clearly delineated by age, location, and symptom status—not the entire population, nor broadly a group with a diagnosed depressive disorder.
Participants were randomized in a one-to-one-to-one ratio: 49 received access to XiaoE, a dialogue system based on cognitive behavioral therapy, 49 to an e-book, and 50 to Xiaoai, a general conversational assistant. The intervention lasted one week. The study is described as a single-blind, three-arm randomized controlled trial. The primary endpoint was the PHQ-9 after one week and again after one month. Working alliance, acceptance, and usability were also included as nonclinical measures.
The result signal is real, but narrow in time
In the intention-to-treat analysis, PHQ-9 scores in the XiaoE group were lower at both measurement time points than in the e-book and general assistant groups. For the one-week time point, the authors report an overall group difference of F(2,136)=17.011, p<0.001, and an effect size of d=0.51. After one month, the difference persisted but was smaller, with F(2,136)=5.477, p=0.005, and d=0.31. In addition, per-protocol analyses, procedures for handling missing data, and sensitivity analyses were conducted.
The randomization supports the interpretation that assignment to XiaoE, under these study conditions, made a difference compared with the two comparison offerings. The finding does not automatically extend further. The PHQ-9 measures the severity of reported symptoms; lower scores neither demonstrate a lasting change nor equivalence to professionally responsible treatment. The last measurement time point was one month after the start. The authors themselves call for longer interventions, longer follow-up, and comparisons with stronger active control conditions. This is not a side note but the decisive limitation of the study’s scope.
Working alliance without mutual responsibility
The result on working alliance is particularly interesting. The XiaoE group rated it better than the comparison groups; the reported group difference was F(2,145)=3.407 at p=0.04. Acceptance was also higher. For usability, however, no statistically significant difference was found. XiaoE was therefore not simply superior on all measured product dimensions. Its distinctiveness lay more in the fact that the interaction was perceived as more fitting and more alliance-like.
But a questionnaire score on working alliance is not evidence of a human relationship. It shows that participants perceived characteristics such as goal orientation, collaboration, or fit in the interaction, to the extent that the instrument used captures these. The system does not thereby assume mutual responsibility or develop its own understanding of the person. Above all, the study does not demonstrate that the better alliance caused the lower PHQ-9 score. Both findings occur together; a test of this mechanism is not reported in the summary. Perceived quality must not be relabeled here as a proven cause of effect.
The conversation format could work—but through what?
The three study arms allow for a cautious but not unambiguous interpretation. The e-book offered a different form of delivery, while Xiaoai also enabled a conversation, but without the specific cognitive-behavioral orientation of XiaoE. That XiaoE performed better than both argues against the simple explanation that either mere information or any kind of dialogue would suffice. What remains unresolved, however, is which component made the difference: the content, the conversation structure, the regularity of use, the fit with the target group, or the interplay of these elements.
Product development in particular tends to mistake the visible surface for the active ingredient. In dialogic software, that surface is the fluent conversation. The XiaoE study suggests a more sober reading: conversational form and professionally structured content can together constitute an effective short-term offering. But it does not provide a decomposition of the individual components. For research, such a question of mechanism would be more important than the claim that the system already possesses a digital counterpart to human empathy.
Engagement Is More Than Decoration and Less Than Efficacy
The pilot study by Kien Hoa Ly, Ann-Marie Ly, and Gerhard Andersson extends this point. It examined a fully automated conversational agent for promoting psychological well-being over two weeks and combined a randomized design with qualitative methods. Participants opened the application an average of 17.71 times during the study period. The qualitative data also brought forth a theme that the authors described as a moderating format. The present summary, however, does not define this feature more precisely, so any further interpretation would be speculative.
The pilot trial primarily supports the assumption that dialogic formats can achieve considerable usage. It does not prove that frequent opening by itself produces psychological improvements. Rather, the authors viewed the good adherence as a reason for a replication with a larger sample and an active control group. In this lies an important methodological discipline: engagement is a prerequisite for a digital intervention to be able to take effect at all. But it is neither a clinical endpoint nor a substitute for a robust comparison.
Large Language Models Are a Different Subject
Zhang and Wang extend the discussion to general language models such as ChatGPT. They cite, among other things, studies in which such models were able to linguistically recognize and formulate emotional components of hypothetical situations. At the same time, they make clear that this performance is based on pattern recognition and language modeling, not on experienced emotions. A model can therefore produce an empathy-sounding response without feeling compassion. Moreover, a test of linguistic emotion depiction is not a clinical efficacy study.
This distinction is central to the comparison with XiaoE. The randomized study tests a specifically oriented system in a defined intervention. The 2024 contribution, by contrast, discusses a technologically and institutionally much broader future. Its references to short follow-up periods, small study groups, limited long-term memory, algorithmic biases, and necessary human oversight relativize its own far-reaching visions of the future. The capabilities of general models can no more be used to derive the efficacy of a specific offering than XiaoE can be used to infer the suitability of every language-generating system.
The unanswered questions lie outside the chat window
The XiaoE results come from a single university, from the COVID-19 pandemic era, and from a very brief intervention. This does not automatically undermine the internal comparison of the randomized groups, but it does limit generalizability. Based on the provided sources, it is not possible to assess whether comparable results would occur in older adults, in other cultural and institutional contexts, with longer use, or with more pronounced and complex complaints. Likewise, the available information does not provide a robust answer regarding the sustainability of the effect.
Data protection, autonomy, crisis situations, and continuity across many conversations are also not clarified by this trial. Zhang and Wang name such questions, but they do not yet provide an empirical solution. Our editorial conclusion is therefore deliberately narrow: XiaoE provides credible evidence of a short-term difference in young adults with depressive symptoms and of a more positive assessment of the working alliance. It provides no evidence of a digital relationship partner or of the replacement of professional treatment. Anyone who makes more of the measured value not only oversteps the data, but also confuses compelling conversation design with responsible care.
Sources & further reading
- Yuhao He, Li Yang, Xiaokun Zhu (2022): Mental Health Chatbot for Young Adults With Depressive Symptoms During the COVID-19 Pandemic: Single-Blind, Three-Arm Randomized Controlled Trial
- Kien Hoa Ly, Ann-Marie Ly, Gerhard Andersson (2017): A fully automated conversational agent for promoting mental well-being: A pilot RCT using mixed methods
- Zhihui Zhang, Jing Wang (2024): Can AI replace psychotherapists? Exploring the future of mental health care