Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
A test for the moment when chatting is no longer enough
The main study, published in Scientific Reports in 2025, selected its test systems through application directories. Included were offerings whose descriptions promised benefits for psychological distress and that contained an AI-based conversation function. In total, the authors examined 29 such agents. This was not an investigation of arbitrary general-purpose models, but of systems offered in the context of mental health. It is precisely this selection that makes the test relevant: people encounter these products in a context in which indications of suicidality can in principle occur.
The researchers confronted all agents with a standardized set of inputs based on the Columbia-Suicide Severity Rating Scale that simulated an increasing suicidal risk. The responses were evaluated according to predefined criteria. These included the ability to provide emergency contacts; the study also considered further factors that are not individually listed in the available source references. The design thus tested not general conversation quality, but a specific safety performance under controlled conditions.
No system passed the original standard
The central finding leaves little room for flattering interpretations. None of the 29 tested agents met the original criteria for an appropriate response. Only under relaxed requirements did 51.72 percent reach the category of a merely borderline response; 48.28 percent were still rated as inappropriate. Common errors were missing emergency information and insufficient understanding of the conversation context. The study thus describes not a rare marginal failure of individual products, but a problem distributed across the sample.
The relaxed criteria deserve particular attention. They show that roughly half of the systems produced at least parts of an expected response. That is better than complete unresponsiveness, but it is not equivalent to robust crisis management. From an editorial perspective, it would be wrong to reinterpret this intermediate category either as a success or to dismiss it as entirely worthless. Rather, it marks the gap between a response that contains individual safety elements and behavior that meets a predefined standard of appropriateness. It is precisely this gap that is decisive for products in sensitive contexts.
A Simulation Test Proves Neither Harm nor Protection
The strength of the design lies in its comparability: All systems received standardized inputs ordered by increasing risk and were assessed according to the same criteria. This makes it possible to show whether elementary response patterns were available under controlled conditions. However, the study does not allow any statement about how often real users experience such situations, how they react to individual responses, or whether a specific response would have caused or prevented harm. Simulated behavior is not a measured care outcome.
The opposite overextension would also be inadmissible. Passing the test does not allow one to derive a protective effect in everyday life. People do not necessarily formulate crises in a standardized form, and conversation trajectories are not identical to individual test inputs. However, this limitation does not arbitrarily weaken the negative finding: Those who fail to provide central information or miss the context even in controlled scenarios have not passed a fundamental functional test. The study does not establish a real harm record, but it does establish a considerable gap in the tested responsiveness.
Recognizing and Responding Appropriately Are Different Tasks
The systematic review by Zhijun Guo and colleagues places the finding within broader research on large language models. It captured 40 English-language articles from the period from January 1, 2017, to April 30, 2024, searched in five databases according to PRISMA guidelines. 15 studies dealt with the detection of mental illnesses or suicidal thoughts through text analysis, seven with language-model-based conversational agents, and 18 with further applications and evaluations in the field of mental health.
The review does report constructive results: Large language models showed good performance in detecting mental health problems and could support easily accessible, less stigmatizing digital offerings. At the same time, it identifies inconsistent outputs, fabricated content, lack of interpretability, data protection issues, and the absence of a comprehensive, comparable ethical framework. For the guiding question of this article, the separation of tasks is particularly important. A model can successfully classify linguistic signals and still fail to generate an appropriate crisis response. Detection, conversation, and crisis management are not the same task, technically or evaluatively.
Earlier Research Explains the Seductive Shortcut
As early as 2020, the systematic review by Aziliz Le Glaz and co-authors showed how widely machine learning and natural language processing had been used in mental health. Of 327 identified articles, 58 were included in the qualitative analysis. The studies worked primarily with medical records and social media posts. Their goals included extracting symptoms, classifying severity levels, comparing treatment outcomes, and deriving psychopathological clues.
These methods can extract information from data that is otherwise hard to access for care, such as documented everyday habits. However, the review also identifies heterogeneous methods, imprecise cohorts in social media, and a preference for powerful rather than transparently functioning classifiers. Moreover, the examined methods often confirmed existing clinical hypotheses instead of producing entirely new insights. This is not a null finding. Rather, it means that statistical strength in a clearly delimited analysis task does not automatically establish comprehensive conversational competence. It is precisely this fallacy that shapes many expectations of today’s generative systems.
The product promise extends further than the tested role
The lead study examined offerings that claimed to provide benefit for psychological distress; it does not follow that each of them explicitly promised crisis assistance. Nevertheless, a practical role problem arises. A conversational system cannot determine that users speak only within the intended product category. Those who explore distress in dialogue may also encounter statements that indicate increasing suicidality. The boundary between general support and a crisis situation runs through the actual conversation, not neatly along a description in the application directory.
Our editorial position is therefore narrower than the blanket demand that every digital offering must resolve every psychological emergency. That would be neither derivable from the sources nor a sensible standard. However, a system must be able to recognize and address the point at which its ordinary role no longer holds. Addressing here does not initially mean taking over the entire crisis itself. Rather, the tested minimum scope includes a context-appropriate response and the availability of relevant emergency information. The fact that many agents already failed at this turns an abstract role question into a concrete product problem.
Crisis capability must become visible as a distinct function
The three publications do not yield an overall judgment about all language-based applications in mental health. The reviews show serious potential in text analysis, early detection, accessibility, and clinical support. At the same time, the study by Pichowicz and colleagues shows that these potentials do not replace a separate examination of crisis responses. Anyone who treats high classification performance, fluent language, or a positive user experience as indirect evidence of safety mixes different target variables. Perceived quality is no more crisis competence than correct recognition is already appropriate action.
For research and product development, the main implication is a more precise language about capabilities. Instead of attributing competence for mental health to a system in a blanket manner, it should be recognizable in each case which task was examined: identifying signals in texts, dialogically accompanying distress, providing information, or reacting to increasing risk. This is not merely a question of documentation. Only separate performance concepts prevent success in one task from legitimizing the next untested task. The lead study does not provide a complete evaluation standard for this, but it does provide a clear reason to treat crisis responses as an independent product property.
The red line runs at the transition to the emergency
The three sources therefore do not tell a simple story of artificial intelligence failing in a human domain. The two reviews document useful analytical achievements and potential supportive roles. Precisely for this reason, the crisis test is so revealing: it separates limited, demonstrable benefit from a more far-reaching assumption about reliability. A system can be helpful in many tasks and still remain inadequate for a specific safety-critical transition. Agreement with its usefulness and criticism of its crisis performance do not contradict each other.
The most important finding of the study, in our view, is not that machines do not replace human professionalism. This general juxtaposition would be coarser than the problem under investigation. What is decisive is that none of the 29 offerings met the original appropriateness standard, even though all occurred in an environment of psychological stress. Thus the debate shifts from impressive answers to verifiable transitions: Does the system recognize that the normal conversation logic ends, and does it then react according to a robust pattern? As long as this question remains unanswered or is answered in the negative, linguistic sovereignty must not be regarded as a safety feature.
Sources & further reading
- W. Pichowicz, Marian Kotas, Patryk Piotrowski (2025): Performance of mental health chatbot agents in detecting and managing suicidal ideation
- Zhijun Guo, Alvina G. Lai, Johan H. Thygesen (2024): Large Language Models for Mental Health Applications: Systematic Review
- Aziliz Le Glaz, Yannis Haralambous, Deok-Hee Kim-Dufor et al. (2020): Machine Learning and Natural Language Processing in Mental Health: Systematic Review