Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

Article history

First published

A randomized pilot with three distinct evaluation tasks

Klos and colleagues studied Tess in 181 Argentine students aged 18 to 33; 158, or 87.2 percent, of the participants were women. After randomization, 99 individuals received access to Tess for eight weeks, while 82 individuals received a psychoeducation book about depression. The research question was deliberately broader than a pure efficacy test: feasibility, acceptance, and a possible influence on depressive and anxious symptoms were to be examined. These three levels are related but methodologically not interchangeable.

The design has a decisive advantage: Tess was not only examined before and after use, but compared with a control condition. This at least makes it possible to test whether the development in the Tess group differed from that in another low-threshold information resource. For depressive symptoms, the research team used nonparametric comparisons, and for anxiety symptoms, t-tests. Because the publication was explicitly designed as a pilot study, it should not be read as a conclusive efficacy test nor as a mere product test.

In the end, 73 evaluable participants remained

The greatest limitation is already visible in the response rate. After eight weeks, data were available from 39 of the 99 individuals in the Tess group and from 34 of the 82 individuals in the control group. This corresponds to 39 and 41 percent, respectively. The substantial loss thus occurred to a similar extent in both conditions. Similar rates, however, do not prevent the remaining individuals from differing from those who did not provide final data.

For the interpretation, this means: randomization was the right starting point, but the final comparison is based only on a portion of the original sample. The source provides no basis for determining motives for dropout or attributing specific experiences to the missing individuals. One must therefore neither claim that they rejected Tess nor assume that their symptom trajectories corresponded to those observed. Additionally, the predominantly female sample and the narrow context of Argentine universities limit the transferability to other populations and care situations.

The group comparison is sobering

From baseline to the eighth week, the researchers found no significant difference between the Tess group and the control group in either depressive or anxious symptoms. For the central question of whether the conversational system has an advantage over the psychoeducation book, this is the decisive finding. The study thus does not demonstrate that Tess reduced symptoms more than the comparison condition. Nor does a non-significant difference demonstrate the equivalence of the two offerings or their lack of efficacy; such a conclusion would require more than this pilot can deliver.

Within the Tess group, anxiety symptoms decreased significantly, while no corresponding significant change was observed in the control group. For depressive symptoms, no significant changes emerged in either group. This is precisely where statistical discipline is needed: the fact that a change is significant in one group and not in the other does not prove a significant difference between the groups. The authors therefore plausibly speak of a possible influence and promising initial results, not of proven superiority.

472 messages are a behavioral signal

In the Tess condition, an average of 472 messages were exchanged, with a standard deviation of 249.52. On average, 116 messages came from users as responses to the system. These numbers say nothing about whether each message was substantive or triggered a psychological change. They do show, however, that for the remaining participants, the application consisted not merely of a one-time opening or a brief trial. Interaction took place to a considerable extent.

A higher number of exchanged messages was also associated with more positive feedback. This correlation is also interesting, but cannot be causally resolved. Perhaps more strongly engaged individuals rated Tess more positively; perhaps a good experience led to more use; possibly both processes worked together. The pilot does not answer this. For product teams, the finding remains valuable nonetheless: a text-based psychological offering can generate repeated engagement. However, engagement is a usage metric and not a surrogate measure for symptom improvement.

Wysa amplifies the engagement signal and its interpretation problem

The earlier Wysa study by Inkster, Sarda, and Subramanian expands this picture with anonymous usage data from a globally available app. The study examined individuals with self-reported depressive symptoms who had voluntarily installed Wysa, used it in a text-based manner, and completed the PHQ-9 at two consecutive time points. Based on usage intensity, groups with high and low usage emerged. Among the 108 heavy users, the self-reported depression score decreased by an average of 5.84 points; among the 21 light users, it decreased by 3.52 points; the difference was statistically significant.

Furthermore, 67.7 percent of the submitted feedback described the experience as helpful and encouraging. This is a serious positive finding regarding user experience. However, because the groups were not randomly assigned to high or low usage, it remains open whether more intensive use caused the greater improvement. Engagement could equally be an expression of motivation, fit, or other differences among users. In comparison, the Tess pilot therefore reveals something important: a randomized design can provide the more rigorous answer to the question of efficacy, even when results appear weaker.

Acceptance is not a consolation prize

It would be wrong to dismiss the Tess study as inconclusive because of the lack of a group difference. Psychological digital products cannot have any supportive effect if people do not use them at all. Repeated conversations and positive feedback are therefore independent developmental findings in their own right. They can indicate that the format, accessibility, or conversational mode resonated with part of the target group. The study thus provides plausible indications of usability and acceptance in the context it examined, exactly as the authors emphasize in their conclusion.

The limitation, however, must be mentioned in the same breath: acceptance is visible here primarily for those individuals who provided data until the end. The high dropout rate prevents a simple narrative of general enthusiasm. Our editorial assessment is therefore neither disparaging nor euphoric. An offering that is actually used has a practical advantage over an unused offering. Whether this advantage leads to reliably better psychological development is a different research question.

The replacement question is larger than the present comparisons

In 2024, Zhang and Wang discuss a far more comprehensive question: whether AI could take over tasks from psychotherapists or even replace them. Their contribution categorizes applications from prediction and monitoring to digital interventions and points to potential advantages in accessibility, scaling, and continuous availability. At the same time, it mentions the lack of genuine empathy, limited long-term memory, algorithmic biases, data protection and autonomy issues, as well as the necessity of human oversight. Regarding long-term effects, the authors describe the state of research as uncertain and call for further large randomized trials.

This debate, however, must not be conflated with the two application studies. Tess was tested against a psychoeducation book, not against psychotherapists. In the Wysa study, heavy and light users of a voluntarily installed app were compared, again not professional practitioners. Also, tests in which a language model recognizes emotions in hypothetical texts or responds to them in linguistically appropriate ways do not measure the same performance as care provided responsibly over time. Among the three sources, there is therefore no direct evidence that an AI system replaces human psychotherapy or equals it.

The appropriate role begins between book and consultation hour

The Tess pilot suggests a more sober role than the grand replacement debate. Its real point of comparison was a book: on one side static psychoeducation, on the other an interactive, text-based system. The concrete product value may lie in this in-between space. A conversational format responds, structures attention, and invites repeated engagement. The fact that students exchanged many messages argues for taking this difference seriously. The fact that symptom trajectories did not withstand the group comparison also limits what can be promised on that basis.

For Dialogatlas, this yields a clear position: Psychological AI should initially be described by its demonstrated function, not by the largest possible professional role. In Tess's case, this function so far is mainly accepted, repeatedly used conversational support in a limited university context. The open question of efficacy does not diminish this finding; it only prevents its overextension. Progress would not consist in passing off usage as treatment success, but in continuing to examine engagement and symptom effect separately. It is precisely this separation that makes the pilot professionally productive despite its sober group comparison.

Sources & further reading