Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

The field is young – and looks older than it is

Kolding and colleagues conducted their review according to the PRISMA guidelines. They searched three databases, had the hits independently checked by two researchers, and qualitatively evaluated content, themes, and findings. In the end, 40 studies were included. The median publication year was 2023. This number alone puts the supposedly long debate into perspective: a large part of the literature emerged only shortly after the broad public availability of generative language models.

The research focused primarily on mental health and well-being in general. Individual disorders played a smaller role; among them, substance use was most frequently addressed. Language-generating models, especially ChatGPT, were at the center. This fits public perception but limits transferability. A finding about a specific model version, a particular prompt, or a simulated situation is not timeless knowledge about generative AI.

The field appears more mature because the systems’ responses seem linguistically sophisticated. However, scientific maturity does not arise from fluent output. It arises from repeatable methods, relevant comparison conditions, real usage situations, and a traceable connection between test and claimed benefit.

A prompt experiment answers a narrow question

The majority of the included studies were prompt experiments. In addition, there were surveys, pilot studies, and case reports. In a prompt experiment, a model receives one or more tasks, and its responses are evaluated against predefined criteria. This can be useful for identifying types of errors, comparing formulations, or checking whether a model can reproduce specialist information in principle.

Such an experiment, however, primarily examines the response to the task posed. It does not automatically capture whether people use the system meaningfully over a longer period, whether corrections are retained, how trust develops, or how a product handles unexpected situations. The user interface, data storage, model changes, moderation, and users’ expectations are also often outside the scope of the test.

Our assessment is therefore not that prompt experiments are mere playthings. They are an early building block. The problem arises when a good comparison of responses is immediately turned into a promise of care provision. Then the narrow research question is retrospectively made larger than the design allows.

Psychiatry is linguistic – but not only text

The authors justify the particular relevance of generative AI with the central role of language in psychiatric diagnostics and psychotherapy. That is convincing. Complaints, relationships, biographical experiences, and changes are to a considerable extent explored through language. A system that processes language flexibly can therefore take on tasks that older rule-based programs could only handle to a limited extent.

Nevertheless, psychiatric and psychotherapeutic work does not consist of generating appropriate sentences. Meaning arises in the course, in knowledge of a specific person, in nonverbal signals, in institutional responsibility, and in decisions that have consequences. A model can produce a linguistically plausible response without possessing that relationship or responsibility.

Precisely because text is so important, one must not underestimate its persuasive power. Good language can make competence visible, but it can also feign competence. The right research question is therefore not only whether a model can respond appropriately, but under what conditions people take its response to be more than it actually is.

The older NLP research had different tasks

The systematic review by Le Glaz and colleagues from 2020 examined machine learning and natural language processing in mental health before today's boom of generative models. It identified 327 studies and included 58. Common data sources were medical documentation and social media. The tasks included, among others, symptom extraction, severity grading, comparison of therapy outcomes, and the extraction of psychopathological clues.

This research was technically less spectacular, but in one respect often clearer: the system had a defined analysis task. Its performance could be measured by a classification or information extraction. The review also found that efficient classifiers were often preferred over more transparent methods, that social media represent an imprecise population, and that language-specific features hinder transferability.

Generative AI expands the possibilities but easily blurs the task. An open conversation can inform, motivate, organize, calm, or advise. Without a clear definition of the goal, almost any plausible answer can be presented as a success. The older research therefore reminds us of a simple rule: before quality is measured, it must be clear what performance is actually expected.

Good Performance in a Test Is a Start

Kolding and colleagues report that generative AI often performed well in the evaluated literature. At the same time, the studies raise significant safety and ethical questions. Both can be true simultaneously. A model can provide useful answers in a defined test and still be insufficiently safeguarded as a product.

In product development, a good test result should therefore not be treated as an endpoint. It is an invitation to the next, more difficult examination. Does the performance hold up with unclear inputs? How does the system react to contradiction? Does it stay on task? What happens when a person understands the answer differently than intended? Which data are processed, and who notices recurring errors?

These questions do not make a project unnecessarily complicated. Rather, they translate the language model performance into a concrete usage situation. Whoever skips this step sells a capability of the model as a property of the entire offering.

Transparency Begins with the Method

The overview calls for more transparent methods, experimental designs, clinical relevance, and involvement of users or patients already in the design phase for future research. These demands hit a sore spot. With rapidly changing models, specifying a brand name is not enough. Version, timing, prompt, settings, and selection of examples influence the result.

Methodological transparency is also important because negative results and unspectacular errors are less often publicly visible than impressive dialogues. A selected example can show that a good answer is possible. For reliability, however, what matters is how often relevant errors occur and under which conditions they accumulate.

Dialogatlas therefore considers public error categories and unknown test trajectories to be more important than a gallery of successful answers. This is an editorial position, not a conclusion of the three sources. But it follows from the gap that the systematic review describes: a young field needs comprehensible procedures more urgently than even more claims about its potential.

Users are not a later validation group

The demand to involve users and patients in design is more than a friendly participation note. Experts and developers often define success through correct, safe, or guideline-compliant answers. Users additionally experience pace, tone, repetitions, paternalism, misunderstandings, and the handling of a no. These characteristics determine whether a system is experienced as helpful in everyday life at all.

Participation does not mean implementing every preference without scrutiny. People may like an answer that is professionally problematic, or experience a necessary boundary as disruptive. But without their perspective, the assumptions the product makes about the conversation remain invisible. Especially in psychological distress, the form of interaction can be as relevant as the information content.

A viable system therefore requires multiple levels of testing: professional and technical tests, real user experience, observation of longer trajectories, and a clear decision about which errors weigh particularly heavily. No single level replaces the other.

From model testing to a responsible offering

The 40 studies in the lead publication show a vibrant and rapidly growing research field. They demonstrate that generative models are being seriously investigated in psychiatric and mental health contexts and can perform convincingly on certain tasks. They do not demonstrate that a general care model has already emerged.

Zhang and Wang describe plausible roles for AI in access, support, data analysis, and personalized offerings, but also emphasize limited long-term findings, bias, ethical questions, and human oversight. Together with the older NLP review, a sober picture emerges: the technology expands the toolbox, but each application requires its own task description and its own evidence.

Progress therefore does not lie in underestimating prompt experiments. It lies in assessing them correctly. A good answer shows one possibility. A good product must turn that into a repeatable, transparent, and accountable performance. Only then does the question arise of what place generative AI should actually take in mental health.

Sources & further reading