Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
A conversation must be continued before it can help
The lead publication from 2017 examined a fully automated conversational agent for promoting psychological well-being. Its focus is not solely on possible changes among participants. What is particularly striking is the reported participation during the two-week intervention: the provided abstract cites a mean of 17.71 app openings over the entire period. The authors rate the adherence as so encouraging that they consider a larger replication study worthwhile.
Our editorial thesis is: return should count as an achievement in its own right for psychological conversational software, not merely as a technical operational metric. An automated offering that is abandoned after first contact offers little opportunity for its content to take effect. Conversely, frequent opening proves neither an improvement in well-being nor a favorable effect on depressive symptoms. Engagement initially describes a relationship between offering and usage, not between offering and health outcome.
The pilot study examines more than a single endpoint
Ly, Ly, and Andersson presented a randomized pilot study using a mixed-methods approach. In doing so, they combined controlled allocation with qualitative data. The research question was directed at a fully automated conversational agent for promoting mental well-being. The study period was two weeks. The combination makes particular sense for an early pilot: quantitative measures can capture differences and usage, while qualitative statements can provide indications of how the format was experienced.
The provided source material, however, does not specify the number of participants, the exact control condition, or the complete outcome values of the well-being measures examined. It is therefore not possible to reconstruct a reliable statement about the size and certainty of a possible effect. What is reliably established here are the high usage, the qualitative subthemes, and the authors’ conclusion that future replication should use a larger sample and an active control group. This limitation is important: the pilot justifies further examination, not yet a general judgment of efficacy.
The moderating format is an unusual clue
In the qualitative data, the researchers identified a subtheme that they described as the moderating format of the conversational agent. In the present material, the term does not denote a statistical moderator. Rather, it refers to an experienced property of the dialogue format. Exactly how this moderation worked, which conversational elements were responsible for it, and how frequently the theme occurred cannot be inferred from the provided abstract excerpt.
Nevertheless, the finding opens a productive perspective. It is plausible that not only the content of a digital intervention matters, but also the way in which questions, answers, and next steps are organized sequentially. This is an interpretation, not a proven mechanism. For product development, it is nonetheless more substantive than the blanket assumption that maximally human-like language automatically creates engagement. Perhaps it is precisely a recognizable, structuring conversational role that contributes to return visits.
17.71 openings are precise and at the same time ambiguous
The number is vivid, but it should not be overinterpreted. An app opening does not indicate how long or intense a conversation was, which contents were addressed, or whether a session was experienced as helpful. The mean also conceals possible differences between individual usage trajectories. The source material provides no distribution for this. The metric therefore shows exposure and interest, but not automatically conversation quality.
The authors assess the usage as higher than in earlier studies of other fully automated offerings, including Woebot and Panoply, which were also described as particularly engaging. Such a comparison provides context, but it is not a direct competition under identical conditions. Different target groups, time periods, and measurement methods may play a role. Editorially, therefore, it is not the ranking that convinces us, but the methodological impulse: adherence is made visible instead of being silently assumed for digital offerings.
Wysa shows the promise and the problem of real usage
The Wysa study by Becky Inkster, Shubhankar Sarda, and Vinod Subramanian supplements the pilot finding with anonymous global usage data. The observed individuals had voluntarily installed the app, interacted with it via text, and reported depressive symptoms on the PHQ-9 at two consecutive assessments. Based on the extent of use, two groups emerged: 108 heavy users and 21 light users. The average improvement was 5.84 and 3.52 points, respectively; the difference was statistically significant, with a reported effect size of 0.63.
This is an encouraging association, but not a causal effect estimate. The groups were not randomly assigned to different usage levels. People who returned more frequently may also have differed in motivation, baseline status, or other unreported characteristics; likewise, an improvement noticed early on may have encouraged further use. In addition, there is a positive experience signal: 67.7 percent of the feedback submitted described the app as helpful and encouraging. This judgment documents acceptance, but it neither replaces a control group nor an independent comparison of outcomes.
XiaoE separates conversation content from mere access
The 2022 study on XiaoE sets up a more sophisticated comparison. In the single-blind, three-arm randomized study, 148 young adults with depressive symptoms at a Chinese university were assigned to a one-week intervention: XiaoE with cognitive-behavioral therapy–based content, an e-book, or the general conversational agent Xiaoai. The PHQ-9 was administered after one week and again one month later. The analysis included intention-to-treat and per-protocol analyses as well as methods for handling missing data.
In the intention-to-treat analysis, the PHQ-9 scores of the XiaoE group were lower than those of the two comparison groups at both time points. Effect sizes of 0.51 after one week and 0.31 after one month were reported. Working alliance and acceptance also differed in favor of XiaoE, while no significant group difference was found for usability. This suggests that access to a dialogue window alone was not decisive. Nevertheless, a short intervention, a single university, and the limited follow-up period remain clear limitations in scope.
Three studies answer three different questions
The publications cannot be meaningfully merged into a single body of evidence. The pilot by Ly and colleagues makes adherence and the experienced conversation format visible, but based on the material available here, it cannot support a precise estimate of effect. The Wysa data show under real-world conditions that intensive use was associated with greater symptom improvement. XiaoE, through randomization and two comparison offerings, provides the strongest basis for the claim that a specifically designed program can make a difference.
The nonclinical measures also deserve a sober reading. A higher-rated working alliance means that participants reported certain qualities of collaboration; it does not prove a human relationship. Consistent usability alongside differing acceptance simultaneously shows that technical operability and content fit are not the same thing. Perceived quality, repeated use, and symptom change can occur together, but each must be examined with appropriate instruments.
The decisive product question is not only: Did it work?
From our perspective, the strength of the lead publication lies in taking seriously an early, unspectacular part of digital impact. A conversational offering must enable a rhythm in which people return and repeatedly engage with content. Those who consider only a final outcome may overlook why a product was used at all. Those who instead optimize only openings, conversation lengths, or positive feedback confuse attention with benefit.
The three sources therefore do not yield blanket agreement with psychological conversational software, but also no reason to downplay their positive findings. Return is a valid development goal, provided it is described as an interim finding. An active comparison subsequently tests whether more occurs than mere engagement with an app. Longer observation must show whether differences persist. Good research begins here with a linguistic discipline that also applies to good products: use is use, experience is experience, and demonstrated change is a third thing.
Sources & further reading
- Kien Hoa Ly, Ann-Marie Ly, Gerhard Andersson (2017): A fully automated conversational agent for promoting mental well-being: A pilot RCT using mixed methods
- Becky Inkster, Shubhankar Sarda, Vinod Subramanian (2018): An Empathy-Driven, Conversational Artificial Intelligence Agent (Wysa) for Digital Mental Well-Being: Real-World Data Evaluation Mixed-Methods Study
- Yuhao He, Li Yang, Xiaokun Zhu (2022): Mental Health Chatbot for Young Adults With Depressive Symptoms During the COVID-19 Pandemic: Single-Blind, Three-Arm Randomized Controlled Trial