Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

The decisive question concerns the relationship between use and benefit.

Inkster, Sarda, and Subramanian did not want to assess Wysa solely by whether people open the application or rate it favorably. Their preliminary real-world evaluation examined whether changes in the Patient Health Questionnaire-9, or PHQ-9 for short, could be observed among users with self-reported depressive symptoms and whether these changes were related to usage intensity. In addition, the team analyzed feedback within the app. This combination of symptom trajectory, actual use, and qualitative experience is essential to the guiding question: A conversational system can only provide support if people actually engage with it.

This is where a robust advantage of the study lies. It does not examine a mere demonstration of the system or a hypothetical conversation, but behavior in ongoing app operation. In doing so, it captures a dimension that controlled trials can easily underestimate: whether an automated conversation generates sufficient engagement across multiple contacts. The flip side follows immediately. Because use was not experimentally assigned, precisely the engagement that makes Wysa interesting is also the central uncertainty in interpretation.

A field observation without random assignment

Anonymous users from various countries were observed who had voluntarily installed Wysa, communicated with the application via text, and reported depressive symptoms using the PHQ-9. For the quantitative analysis, two groups were formed based on use at and between two consecutive screening time points: 108 people with high use and 21 with low use. These groups were thus a result of the observed behavior. No one was randomly assigned to write particularly much or little with Wysa.

This construction is more important for understanding than the label artificial intelligence. Existing usage trajectories were compared, not two interventions administered under otherwise identical conditions. The study can therefore examine a correlation between more intensive use and greater improvement. It cannot isolate whether the intensity of use itself caused the change. The comparison of unequally sized groups also complicates a simple reading. That does not make the study worthless; rather, it precisely defines what kind of statement its data can support.

The numbers are encouraging, but not self-explanatory

In the high-use group, the mean improvement in self-reported depression scores was 5.84 points, with a standard deviation of 6.66. In the low-use group, it was 3.52 points, with a standard deviation of 6.15. The group difference was statistically significant in the Mann-Whitney test used, with P equal to 0.03; the authors report a moderate effect size of 0.63. The study thus contains a genuine positive signal: stronger use was associated not only with activity in the app, but also with a greater mean improvement in the measured values.

Equally important is the wide dispersion within both groups. The averages do not describe a uniform experience for all participants. Moreover, from the provided study findings it cannot be deduced what proportion achieved a specific clinically relevant change or whether the differences persisted longer. The correct statement is therefore neither that Wysa demonstrably works nor that the data mean nothing. What is evidenced is a statistical association in this observed selection of users; the causal explanation remains open.

Engagement can be a cause, a consequence, or a selection mechanism

Several interpretations are plausible. More frequent writing could have led to users receiving more supportive prompts and benefiting more. Likewise, people who noticed an improvement or a fitting conversation experience early on may have continued writing precisely for that reason. It is also conceivable that the groups already differed in motivation, expectations, or other unreported characteristics. These are not proven explanations, but typical alternatives that follow from the observational design.

For product development, this ambiguity is more productive than a hasty judgment of success. A system that is used repeatedly has obviously overcome a relevant hurdle. It reaches people not only technically, but sustains an interaction. However, good engagement does not automatically lead to a specific psychological effect. Our editorial assessment is: engagement in such applications should be taken seriously as an independent outcome, without covertly elevating it to evidence of efficacy. Only when usage dose, baseline, and trajectory are methodologically separated can we more precisely say what the conversation itself contributes.

Positive feedback measures experience, not symptom effect

The qualitative side of the Wysa study strengthens the finding that the application was acceptable to many respondents. Of the feedback submitted within the app, 67.7 percent rated the experience as helpful or encouraging. Precision is crucial here: the figure refers to the feedback responses received, not automatically to all observed users. It describes a perceived quality of the interaction. Whether this perception caused the measured change cannot be concluded from this.

The mixed-methods approach also included the analysis of objections in conversations and the testing of a machine classifier to detect them. However, the present abstract does not provide a concrete result that would allow further assessment. This knowledge gap should not be filled with speculation. Nevertheless, it can be stated: the study did not treat conversation quality solely as mood at the end of a session, but also attempted to systematically expose friction within the interaction.

The randomized pilot sets a different benchmark

The pilot study by Ly, Ly, and Andersson from 2017 extends the comparison because it examined a fully automated conversational agent in a randomized design. The intervention lasted two weeks. High engagement was reported, with an average of 17.71 app openings over the entire period. The qualitative data also provided indications of the system's moderating approach, that is, a conversational function that users apparently perceived as a distinctive feature. The study thus confirms that automated dialogues can indeed generate intensive use over a short period.

The pilot does not, however, conclusively resolve the question of evidence. The authors explicitly justify a repetition with a larger sample and an active control group. Such a comparison condition would be important to better distinguish the specific contribution of the conversational format from general digital engagement or attention. The provided source material lacks information that would allow a detailed assessment of individual outcome measures. The meaningful comparison with Wysa therefore lies primarily in the design: randomization improves causal testing, while real-world data remain closer to actual use.

Large language models do not retroactively change the Wysa evidence

Zhang and Wang in 2024 sketch a considerably broader picture of artificial intelligence in mental health care. Their contribution addresses, among other things, prediction methods, digital interventions, support for professionals, and continuous monitoring. It points to preliminary indications regarding anxiety and depression symptoms, but at the same time emphasizes small study groups, lack of long-term observation, and the need for large randomized trials. It also discusses algorithmic biases, data protection, limited long-term memory, and the necessity of human oversight. This is a programmatic overview, not a randomized experiment of their own described in the provided text.

Particularly important is the distinction between linguistic performance and emotional experience. The contribution attributes to current models a high ability to recognize emotional patterns in texts and generate appropriate responses, but notes that this is based on pattern processing and not on genuine emotional understanding. These newer capabilities must not be read retroactively into the 2018 Wysa study. A specialized mobile system and a general large language model are not interchangeable, either technically or in terms of their tested usage. Advanced formulations do not replace application-related evidence.

The fair claim is smaller and more interesting

The three publications do not provide a serious basis for the sweeping question of whether AI can replace psychotherapists. They deliver something more concrete: automated conversations can be used over short periods, be experienced positively, and be associated with improved self-reports. For Wysa, this association is visible in real-world operation; the randomized pilot project of another system also shows considerable participation. The later overview also makes clear that short-term findings, technical language competence, and long-term care outcomes are at different levels.

The editorially most convincing reading of the Wysa study is therefore neither euphoric nor defensive. Its strongest contribution is to make engagement visible as part of the research topic. People voluntarily continued the conversation, many responses were positive, and intensive use was associated with more strongly improved values. This justifies further, more rigorous examination and the investigation of complementary forms of application. It does not justify turning a user group into a mechanism of action. Anyone assessing psychological AI conversations should therefore first ask whether acceptance, subjective help, symptom change, or causal efficacy is being claimed. The differences are not linguistic pedantry but the core of the evidence.

Sources & further reading