Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

Progress has at least three meanings

The systematic review by Yining Hua and colleagues examines studies from 2020 to 2024 and distinguishes rule-based, machine-learning, and large-language-model-based systems. This classification reveals an often hidden difference: technical progress, practical usability, and health benefit are distinct things. A system can produce more variable responses without thereby reducing symptoms. It can be used regularly without having a proven effect. And a tightly guided dialogue can be useful despite simple technology.

Precisely for this reason, architecture is not a ranking of care quality. According to the review, rule-based offerings dominated until 2023; in 2024, 45 percent of new studies already involved systems with large language models. This demonstrates a rapid research shift, but not yet a corresponding leap in effect. Editorially, we consider this distinction to be the core of the finding: anyone who recognizes innovation solely in generative language confuses a new conversational interface with an already tested health instrument.

160 studies bring order to a confusing field

Source 1 is a systematic review of 160 studies. Its research question addresses not only the results of individual applications but also the development of the entire field: which technical architectures are being studied, at what assessment stage are the studies, and how robust are the statements about therapeutic benefit? The review thus responds to a fragmented literature in which systems with very different modes of operation appear under similar names and in which the rigor of evaluation varies considerably.

However, the provided source material does not reveal the databases searched, detailed inclusion and exclusion criteria, or a concrete assessment of the risk of bias. These methodological details therefore cannot be assessed here. The summary also does not permit any conclusions about how diagnoses, target groups, or study durations are distributed across the 160 papers. The review thus provides a reliable mapping of the reported types of evaluation, but no blanket estimate of efficacy for psychological conversational systems as a whole.

A testing model against conceptual shortcuts

Hua and colleagues propose three stages. The first stage is basic technical testing: Does the system work as intended, and can its responses be technically validated? This is followed by a feasibility check that primarily examines usage and participation. Only the third stage addresses clinical efficacy, in which symptom changes are tested. The model is unspectacular – and precisely for that reason useful. It prevents a good response rating, high usage, or positive feedback from being silently passed off as a health effect.

The stages are not interchangeable. Without technical reliability, there is no basis for further testing; without actual usage, a potentially effective application can remain practically meaningless. Conversely, intensive usage does not prove symptom reduction. Nor is a study at the third stage automatically high-quality: The review summary only states what an investigation aims at, not that every efficacy study was randomized, sufficiently large, or designed for the long term. The testing model creates order, but it does not replace an assessment of the respective design.

For language models, the evidence gap opens up

Across all architectures, according to the review, 47 percent of studies focused on clinical efficacy. For studies on large language models, however, this proportion was only 16 percent. At the same time, 77 percent of LLM studies were still in early validation. These figures do not indicate proven ineffectiveness. Rather, they show that the fastest-growing class of systems has largely not yet been tested on the question that is particularly important in psychologically stressful contexts: Does the application improve relevant health outcomes?

This is a temporal and methodological mismatch. Generative systems quickly reach research, public use, and product development, while clinical trials inevitably take longer. This delay implies neither a prohibition nor a carte blanche. It demands precise language: A technically impressive model is a technically impressive model. A well-received pilot is a well-received pilot. Only an appropriate efficacy test can support further-reaching claims. This linguistic discipline would already be a considerable advance.

Being rated as empathetic is not being effectively treated

John Torous and Charlotte Blease sharpen this separation. They refer to research with nonclinical samples in which AI could improve text-based support, but whose evaluations focused primarily on perceived empathy rather than clinical outcomes. Perceived empathy is by no means trivial: it can influence acceptance and willingness to engage in conversation. But it initially describes an experience with communication. Whether it changes symptoms, functioning, or care outcomes is a different question and requires a different study design.

Their contribution also broadens the perspective beyond direct conversations. Obvious applications accordingly also lie in documentation, administrative work, training, and routine symptom monitoring. This is not an evasion of the grand vision, but possibly the more realistic direction of development. Torous and Blease also warn against the assumption that mere access to digital resources solves care problems. Decades of self-help offerings, internet-based programs, apps, and telemedicine have already shown that availability alone does not guarantee effective prevention.

The replacement question produces more heat than insight

Zhihui Zhang and Jing Wang explicitly raise the question of whether AI could replace psychotherapeutic professionals. Their contribution assembles possible advantages: constant availability, scalability, consistent procedures, potentially more open disclosures to a machine perceived as non-judgmental, and support in regions with few professionals. It also refers to preliminary findings on anxiety and depression symptoms. At the same time, it names small participant groups, missing long-term observation, and indications that short-term effects need not persist over longer periods.

Particularly instructive is the boundary between demonstration and proof. A hypothetical conversation presented in the contribution shows how a language model could mirror distress and offer coping strategies; but it is not a study result. Even good performance on a scale for linguistic expression of emotional awareness does not prove experienced empathy. Zhang and Wang themselves acknowledge this limitation: the model recognizes and generates patterns; it does not feel. In the end, despite far-reaching expectations for the future, they advocate for supplementation rather than replacement and for human oversight.

The AI label also requires scrutiny

The review by Hua and colleagues also finds discrepancies between marketed claims and the actual architecture. Some interventions labeled as AI-supported were accordingly based on simple rule-based scripts. This is not merely a problem of exaggerated advertising. Without a transparent description of the architecture, it is hard to assess which findings are transferable to other systems. A fixed dialogue script has different error possibilities than a generative model that can produce plausible but incorrect answers.

For large language models, the overview cites incorrect answers, data protection risks, and unverified therapeutic effects as particular ethical problems. Torous and Blease add biases from training data, insufficient protection of sensitive health information, and difficult integration into existing care. Zhang and Wang also point to limitations of long-term memory and possible breaks in conversation continuity. None of these dangers proves that generative support is fundamentally unusable. But it determines what must be investigated before a concrete deployment.

The most meaningful product specification would be a testing level

Our editorial position is therefore deliberately narrower than the usual promises about the future: Psychological AI should not be judged primarily by whether it imitates or replaces human professionals. What matters is the specific task it is intended for and the level at which that task has been tested. A system for psychoeducation, a conversational offering for emotional support, and a tool for symptom monitoring make different claims. The same fluid interface does not imply the same requirements or the same permissible conclusions.

The most important contribution of the review is therefore not a final assessment of benefits or harms. It provides a grammar with which claims can be read more precisely. Architecture, technical performance, usage, perceived quality, and clinical change should be expressed in separate sentences. Product development and research would gain credibility if, alongside the model name, they disclosed which of these levels was actually examined. The current lag in clinical testing is not a judgment on the future of large language models. It is a judgment on how little evidence there is for their present state.

Sources & further reading