Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
The 16 studies are a map, not an efficacy judgment
The scoping review published in 2025 in npj Digital Medicine examines large language models for generative tasks in mental health care. The authors aim to bring together existing applications, model performance, and clinical efficacy, identify gaps in evaluation methods, and derive recommendations for further development. A systematic search yielded 726 distinct publications; only 16 met the inclusion criteria. These works dealt with, among other things, clinical assistance, counseling, supportive conversations, and emotional companionship.
The ratio of 726 hits to 16 included studies is striking, but should not be misinterpreted as a quality rate. The provided summary does not reveal all exclusion reasons nor the individual study designs. The review is also designed as a scoping review: it maps a young and heterogeneous field rather than calculating a common treatment effect. A pooled efficacy estimate is not reported. Its most important contribution therefore lies less in a final judgment on benefit than in the diagnosis of why such a judgment is currently difficult to make.
From language analysis to language production
How much the subject matter has changed is shown by comparison with the systematic review by Le Glaz and colleagues from 2020. It covered research on machine learning and natural language processing in mental health. After a search in four medical databases, 58 of 327 identified articles were included. According to the summary, the analysis was organized qualitatively; a quantitative synthesis of effects is not reported there. The studies examined mainly medical records, emergency department data, and social media posts.
The systems of that time were intended, for example, to extract symptoms from texts, classify severity levels, provide indications of psychopathological phenomena, or compare treatment trajectories. Language processing here largely meant statistically evaluating existing language. The current lead review, by contrast, focuses on systems that themselves generate utterances. This is not a mere technical continuation. A classifier outputs a classification; a generative model enters a situation linguistically. This makes comprehensibility, appropriateness, and interaction quality visible – at the same time, it becomes less clear which of these properties can already count as care-relevant success.
Positive signals deserve precise appreciation
Hua and colleagues find initial promising results in the 16 studies. This should not be downplayed. The fact that language models are being tested in different task areas and show usable performance there at all is a relevant research finding. The range of applications extends from clinical support to emotional support. The 2024 review also reports successes, including in accuracy and accessibility, and describes considerable potential for mental health care.
However, more cannot be reliably derived from the available summaries. Information that would allow a comparison of individual applications is missing: specific effect sizes, duration of observed changes, or a uniform description of the investigated usage contexts are not mentioned. The positive statement is therefore not that generative systems have already proven their suitability for care. It is narrower and yet substantial: In various tasks, outcomes emerge that justify systematic further development. Editorially, we consider this distinction crucial. Early successes are a reason for better research, not for hasty devaluation – but equally not for blanket product promises.
A field with many yardsticks of its own
The central finding of the lead review concerns evaluation. Many studies used non-standardized, self-designed scales. According to the authors’ assessment, this limits both the comparability and the robustness of the results. An application can perform convincingly within one study without it becoming clear how its result relates to another study. Even similar-sounding assessments can capture different things: the performance of the model, the quality of individual outputs, or an effect in the care context.
A specially developed scale is not automatically worthless. For novel tasks, there can be good reasons to first design suitable criteria. The diversity becomes problematic when each study defines success differently and no common point of reference emerges. Then the number of positive individual findings does grow, but the shared knowledge does not grow to the same extent. This is precisely where we see the core conflict: The linguistic flexibility of the models expands the possible applications faster than research clarifies which results are comparable between applications.
Standardization should not be confused with a single universal value. Clinical assistance, counseling, and emotional support are different tasks and do not necessarily require identical metrics. Rather, what is needed is a comprehensible assessment architecture: clearly named task, disclosed procedure, and a separation between model performance and clinical efficacy. The lead review calls for stricter development and evaluation guidelines. However, how a binding standard should look in detail cannot be gleaned from the summary.
Why the preprint comes to 34 papers
The systematic scoping review from 2024 appears at first glance to present a different picture. It identified 313 publications and included 34. Its search took place in November 2023 in PubMed, Web of Science, Google Scholar, arXiv, medRxiv, and PsyArXiv. Included were original works, including non-peer-reviewed ones, without language restrictions, provided they had been disseminated within the specified period, used newer large language models, and directly addressed questions of mental health care. The applications included, among others, diagnostics, supportive conversations, and the promotion of participation.
The numbers 16 and 34 therefore do not contradict each other. The lead review from 2025 explicitly focuses on generative tasks, while the preprint considers a broader spectrum of applications. Search spaces and inclusion logics can also alter the selection. The extent to which the two sets of studies overlap is not evident from the provided information. A plausible interpretation is therefore that these are two differently tailored maps of the same growing field – not two direct replications.
In terms of content, their diagnoses converge anyway. The preprint identifies problems with data availability and data reliability, with the differentiated handling of mental states, and with effective evaluation methods. It sees gaps in clinical applicability and ethical questions and calls for robust datasets, assessment frameworks, and interdisciplinary collaboration. The broader selection thus does not eliminate the measurement problem; it makes visible that it extends beyond purely generative conversational tasks.
Transparency is part of the result
The lead review names a second structural bottleneck: Many works rely on prompt-tuning of proprietary models, particularly from the GPT series. This raises, in the authors’ assessment, questions of transparency and reproducibility. An observed result depends not only on the abstract model family but also on the specific configuration and the inputs used. If crucial components are not accessible or are incompletely documented, other research groups can only partially reproduce the investigation.
This tension is older than the current generation of language models. Already the 2020 review noted that powerful classifiers were preferred over transparently functioning methods. Generative systems exacerbate the problem because their answers themselves become the visible product. Our editorial position is therefore clear: In this field, traceability is not a mere technical addition. It is part of the evidential value of a result. Conversely, it would be wrong to equate openness alone with quality. An accessible model is not yet effective, appropriate, or safe; it can merely be better tested.
Autonomy is a much larger claim
Hua and colleagues conclude that the current evidence does not fully support the use of large language models as standalone interventions. This formulation is narrower than a blanket no. It neither disputes initial useful results nor the possibility of clinical integration. Rather, it marks the gap between a successfully handled subtask and an offering that, without further integration, takes on a more comprehensive role.
The older review accordingly describes machine learning and language processing as potential tools to support clinical practice. The 2024 preprint also recognizes potential but emphasizes persistent gaps in clinical applicability. Taken together, the three publications suggest a task-oriented interpretation: a system can be useful in a clearly defined function, even though its suitability for a standalone care role has not been demonstrated. This is not a linguistic caveat but claims of a different magnitude.
For product development and public communication, this implies, in our assessment, a simple but consequential rule: the named role must not be larger than the tested task. Whoever has studied good performance in clinical assistance has not automatically validated counseling or emotional support. The available sources, however, do not permit a blanket determination of which specific division of tasks would be appropriate in each field of application.
First the task, then the judgment
Research does not first need even more impressive dialogue examples, but rather more closely linked statements about what a result means. A convincing study would, in our view, need to specify in advance which task the model fulfills, in which context the task is tested, and whether the assessment instrument captures the system’s output or a clinical effect. Model, prompt procedures, and scales would need to be documented in such a way that a result can be traced and compared with related work. These requirements arise directly from the deficits described by the reviews.
In doing so, the diversity of applications should be preserved. The goal cannot be to squeeze clinical assistance, counseling, and emotional support into the same measurement format. A common framework should instead indicate where comparability is meaningful and where task-specific criteria are necessary. That would be more than methodological order: it would prevent linguistic fluency from inadvertently serving as a proxy for care success.
The lead review thus provides no reason to define generative systems out of mental health care. It provides a more precise mandate. As long as studies measure with their own rulers and proprietary configurations are only partially reconstructable, even a positive result remains hard to assess. The field does not have too few interesting applications. It has too few shared rules for which conclusion an interesting application actually supports.
Sources & further reading
- Yining Hua, Hongbin Na, Zehan Li (2025): A scoping review of large language models for generative tasks in mental health care
- Aziliz Le Glaz, Yannis Haralambous, Deok-Hee Kim-Dufor et al. (2020): Machine Learning and Natural Language Processing in Mental Health: Systematic Review
- Yining Hua, Fenglin Liu, Kailai Yang et al. (2024): Large Language Models in Mental Health Care: A Systematic Scoping Review (Preprint)