Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
The actual point of contention lies between effect and trust
The lead publication asked whether dialog-capable systems can improve mental health and whether their use is safe. It defined conversational systems as applications that interact with people through spoken, written, or visual language. The background was the global shortage of mental health professionals. This care context explains the interest in technical support, but proves neither its efficacy nor its suitability for specific care tasks.
What is decisive, therefore, is not whether such systems fundamentally possess potential. The review explicitly acknowledges that. What is decisive is which conclusion the existing studies actually support. Our assessment: A positive symptom signal, a pleasant conversation, and the absence of documented harms are three different findings. Anyone who combines them into a general quality judgment turns incomplete evidence into a certainty that is not present in the sources.
Broad search, narrow results base
For the systematic review, the authors searched seven bibliographic databases, including MEDLINE, EMBASE, and PsycINFO, as well as Google Scholar. In addition, checks of backward and forward citation references of included studies and relevant reviews were performed. Two reviewers independently selected studies, extracted the data, and assessed the risk of bias. Depending on suitability, the findings were synthesized narratively or statistically. The review was also registered in the PROSPERO registry.
Of 1,048 literature references found, twelve studies were included. Together they examined eight outcome areas. These numbers are more important for interpretation than the meta-analysis label: a statistical pooling can organize scattered findings, but it does not automatically increase the number of high-quality studies per outcome area. The authors explicitly state that only a few studies were available for individual endpoints and that the included studies had a high risk of bias. The methodological breadth of the search thus meets a limited and vulnerable primary literature.
Some symptoms improved, while other findings remained inconclusive
The review found weak evidence that the investigated systems could improve depression, psychological distress, stress, and acrophobia. For subjective psychological well-being, in contrast, no statistically significant effect was shown on a similarly weak evidence base. For anxiety severity as well as positive and negative affect, the results were contradictory. This distribution alone contradicts the notion of a uniform effect on mental health.
The appropriate reading is therefore not that AI conversations work or do not work. The source evidence is narrower: for some measured endpoints there were positive signals, for another no statistically significant effect, and for others no consistent direction. From this, neither general efficacy nor a ranking relative to other offerings can be derived. Such comparisons would only be permissible if they were supported by corresponding study designs and data; the provided sources offer no basis for such comparisons.
Statistical change is not yet a clinically important change
It is no coincidence that the authors limit their conclusion to potential. They saw insufficient evidence for a definitive judgment. In addition to the high risk of bias, the small number of studies per outcome, and the contradictory findings, they also cite the lack of knowledge about whether observed effects were clinically meaningful. In doing so, they mark a boundary that is easily lost in product communication and public debate.
A statistical difference initially answers a statistical question. Whether a change is relevant, lasting, or meaningful in the care context for those affected does not automatically follow from it. The meta-analysis does not substantiate these further properties. Our editorial position is therefore strict: a provider or institution should not verbally upgrade a measured symptom endpoint to comprehensive mental improvement. The scope of a claim must not be greater than the scope of the examined result.
Absence of reported harm is not a reliable all-clear
The basis for statements on safety is particularly thin. Only two of the twelve included studies assessed it. These investigations concluded that the systems were safe because no adverse events or harms were reported. The meta-analysis, however, does not derive any definitive certainty of safety from this, but calls for further research in order to be able to draw viable conclusions on efficacy and safety.
Here lies the core conflict of the article. The absence of reported harms can have various reasons; which of these played a role in the two studies cannot be inferred from the information provided. Precisely for this reason, any further explanation would be speculation. What can be said with certainty is only this: Two studies without reported harm events constitute a very limited observational basis. From an editorial perspective, it is indefensible to derive a general safety promise for AI conversations in the field of mental health from this.
Positive perception answers a different question
The third source expands the picture with the perspective of users. For this scoping review, eight electronic databases as well as forward and backward citation searches were conducted. Two reviewers independently selected studies and extracted their data. Of 1,072 literature references found, 37 independent studies were included in a thematic analysis. This produced ten themes: usefulness, ease of use, responsiveness, comprehensibility, acceptance, attractiveness, trustworthiness, enjoyment of use, content, and comparisons.
Overall, the review reported predominantly positive perceptions and opinions. At the same time, it identified concrete open problems: The linguistic capabilities would need to better process unexpected inputs, generate high-quality and more variable responses, and align content more strongly with individual treatment recommendations. The authors see a need for personalized conversations to address this. These results are relevant for design and implementation. However, they do not prove that positively rated systems reliably improve complaints or rarely cause harm. Perceived usefulness and demonstrated efficacy remain different categories.
The preprint does not double the evidence
The second source is the preprint version of the later meta-analysis, published in 2019. It states the same research question, the same search paths, 1,048 literature references found, twelve included studies, eight outcome areas, and the same central results. The justification for the cautious final judgment also matches. The preprint thus documents the publication history of the lead study, but does not constitute an independent confirmation by a second research group or an additional data set.
This distinction is more than bibliographic pedantry. Three sources here do not mean three independent bodies of evidence: Two entries belong to the same systematic review, the third examines perceptions instead of clinical efficacy and safety. Taken together, they provide a meaningful contrast between results and user experiences. But they do not answer all questions about current AI systems, individual products, different conditions of use, or long-term consequences. These knowledge limits should remain visible.
Good conversations require separate quality judgments
The sources do not yield a blanket verdict against conversational systems. There is weak evidence of improvements for individual complaints and predominantly positive perceptions. Both deserve further investigation. Equally clear, however, is that subjective agreement does not replace missing safety data, and a statistical effect does not automatically have clinical significance. The most interesting question, therefore, is not whether people enjoy talking to a system, but whether the claimed benefit has been demonstrated under robust conditions and whether the potential harm has actually been investigated.
For research, product development, and professional practice, this yields, in our view, a sober standard: efficacy, clinical relevance, safety, and user experience must be reported separately. A system can be comprehensible and attractive without having proven its health effect. It can show a positive signal on one endpoint without being suitable for other complaints. And it can have no reported harm in a few studies without thereby being generally considered safe. A convincing AI conversation in particular must not blur these distinctions. Its linguistic quality does not increase the evidentiary weight of the underlying research.
Sources & further reading
- Alaa Abd‐Alrazaq, Asma Rababeh, Mohannad Alajlani (2020): Effectiveness and Safety of Using Chatbots to Improve Mental Health: Systematic Review and Meta-Analysis
- Alaa Abd‐Alrazaq, Asma Rababeh, Mohannad Alajlani (2019): Effectiveness and Safety of Using Chatbots to Improve Mental Health: Systematic Review and Meta-Analysis (Preprint)
- Alaa Abd‐Alrazaq, Mohannad Alajlani, Nashva Ali (2020): Perceptions and Opinions of Patients About Mental Health Chatbots: Scoping Review