Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

The grand replacement question is too coarse for the existing research

The lead publication by Guo and colleagues examines applications of large language models for mental health in three broad fields: early screening, digital interventions, and clinical use. This structure alone reveals the fundamental problem of the public debate. A model that recognizes mental-health-relevant cues in texts fulfills a different function than a system that conducts a supportive conversation. Both, in turn, differ from a tool that supports professionals in everyday clinical practice. A good result in one field does not imply suitability for the others.

The replacement question blurs these differences. It invites treating linguistic fluency as a uniform professional competence. This is not supported by the review nor by the two comparison sources. Our assessment is therefore methodological: anyone asking about a complete replacement of a professional role demands from research an answer to a package of tasks, relationships, and responsibilities that the studies examined do not jointly represent. A more sober debate begins with the question of which specific performance was tested and how its success was measured.

Forty studies yield a map, not a uniform efficacy value

Guo and colleagues worked according to the PRISMA guidelines and searched MEDLINE via PubMed, IEEE Xplore, Scopus, JMIR, and the ACM Digital Library. English-language articles from January 1, 2017, to April 30, 2024, were considered. A total of 40 publications were included in the analysis. Of these, 15 dealt with the detection of psychological distress or suicidal thoughts through text analysis, seven with large language models as conversational systems, and 18 with further applications and evaluations in the field of mental health.

This is an insightful overview for a rapidly growing research field, but not a homogeneous collection of clinical efficacy studies. The review brings together various models, data sources, methods, and outcome measures. Accordingly, its overall number must not be read as the sample of a single experiment. In addition, there is the restriction to English-language literature. The authors specifically identify the lack of multilingual, expert-annotated datasets as a problem. What works in one language or dataset is therefore not automatically transferable to other populations, modes of expression, or care contexts.

The most tangible benefit lies in text recognition

One should not downplay a positive finding: The review reports good performance of large language models in detecting mental health problems in texts. This offers a plausible benefit for early screening and the analysis of large volumes of text. Models can process linguistic patterns that are relevant for indications of distress or suicidal thoughts. It is precisely here that their scalability and ability to process extensive data bring their technical strengths to bear.

But detection is not yet a proven improvement in care. A classification result merely indicates how a system categorizes texts; it does not by itself demonstrate that people thereby achieve better long-term health outcomes. Nor does a statistical assignment alone justify a clinical decision. This distinction is not a devaluation of the benefit, but rather its refinement. Editorially, we therefore consider the analysis and indication function to be a considerably more robust starting point than the claim that a language model could assume a comprehensive professional role.

Accessible conversations are a real advantage, but not cumulative evidence

Guo and colleagues also see potential for accessible and less stigmatizing digital offerings. Conversational systems are not bound to office hours and can enable a low-threshold form of interaction. This advantage is practically significant: An offering can be valuable because it is reachable and people find it easier to use. The review thus points to a quality that is easily overlooked in a debate narrowed exclusively to clinical endpoints.

Accessibility, acceptance, and efficacy nevertheless remain distinct concepts. A conversation perceived as pleasant or relieving proves neither correct content nor sustainable effects. The publication by Zhang and Wang extends this point: It refers to preliminary studies with possible short-term improvements in anxiety and depression symptoms, but emphasizes small groups, lack of long-term observation, and uncertain durability. The publication does not describe its own randomized investigation; it is therefore to be regarded as a review that presents an argument, not as independent evidence of efficacy.

Even before generative models, the picture was cautiously positive

The 2019 review by Vaidyam and colleagues offers a useful historical comparison. Their systematic search from June 2018 covered six databases. Of 1466 entries found, eight studies met the inclusion criteria; two additional ones were added through the examination of reference lists. The ten studies in total concerned conversational systems for people with mental illnesses or increased risk, including depression, anxiety, schizophrenia, bipolar disorders, and substance use disorders.

The finding at the time was preliminary, but not merely cautionary. The studies reported potential especially for psychoeducation and self-adherence; satisfaction scores were also consistently high. At the same time, heterogeneous study designs and inconsistent outcome reporting prevented a robust overall conclusion on effectiveness. The continuity up to 2024 is remarkable: The technology has become more linguistically capable, yet the fundamental question of evidence remains similar. Positive user experiences and plausible individual applications contrast with a research landscape whose target variables and testing procedures are not yet sufficiently standardized.

Emotionally appropriate language is easily confused with understanding

Zhang and Wang emphasize that current models can recognize and articulate emotional components of hypothetical situations in language. They refer, among other things, to a study using the Levels of Emotional Awareness Scale, in which ChatGPT’s responses were comparable to or above those of the general population. This is an interesting finding about generated language. The authors themselves, however, make clear that the system does not experience emotions but processes patterns and generates appropriate formulations.

This distinction is central to AI conversations. Linguistic empathy can be pleasant, helpful, or open up users, without human experience behind it. Conversely, the lack of experience does not automatically mean that every supportive formulation would be worthless. Our editorial position lies between these fallacies: What matters is not whether the machine inwardly feels, but whether its concrete response is reliable, situationally appropriate, and tested for the intended purpose. A convincing simulation must not be reinterpreted as evidence of diagnostic judgment, long-term relationship continuity, or clinical efficacy.

Clinical use currently fails less because of tone than because of reliability

The sharpest assessment of the lead publication concerns clinical use. Guo and colleagues conclude that its current risks could outweigh the benefits. Mentioned are inconsistent outputs, fabricated content, limited traceability due to the black-box nature, data protection issues, and the lack of a comprehensive, comparable ethical framework. Added to this is the risk of excessive dependence on the systems – both on the part of patients and physicians.

Zhang and Wang add problems of algorithmic bias and limited long-term continuity. Language models cannot easily integrate previous interactions durably and coherently; additional technical processes are needed for that. The comparison of sources thus suggests an important shift: The decisive weakness is not that AI responses are fundamentally cold or unusable. On the contrary, they can seem remarkably fitting. The problem is that this quality can fluctuate and that fluent language makes errors appear particularly credible. The more consequential the use, the less the impression of a good conversation suffices.

A workable division of roles begins below the grand promise

The systematic review does not categorically reject large language models. It explicitly describes them as potential clinical aids and calls for further research and development. The earlier findings on psychoeducation, adherence, and high satisfaction also argue against a reflexive dismissal. Rather, the material yields a graded perspective: recognizing cues in texts, conveying information accessibly, or supporting specialist work are more narrowly defined tasks than independent diagnostics, treatment, and long-term support.

In our view, this is exactly where the debate should begin. Not every AI conversation has to be measured against the maximal claim of being able to replace an entire profession. But every claimed purpose needs its appropriate test: classification quality for recognition, usability for access, consistent content for information, and robust clinical endpoints for health impact. The three publications do not answer all questions about mental health care with AI. However, they suffice for a clear editorial conclusion: large language models neither become psychotherapists through human-like language nor become meaningless through their limitations. Their sensible place emerges where tasks are defined more narrowly than the promise.

Sources & further reading