Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
The replacement question turns answers into a professional role
A language model can explain, mirror, ask questions, and formulate suggestions. In selected dialogues, these abilities can appear impressive. However, it does not follow that the system takes on the role of a psychotherapeutic professional. A professional role does not consist of a collection of good sentences, but of training, relationship, responsibility, institutional rules, and decisions under uncertainty.
Moore and colleagues explicitly examine the use of language models as a replacement for mental health care providers. To this end, they map therapy guidelines from large medical institutions and derive important characteristics of therapeutic relationships. Subsequently, they test current models, including GPT-4o, in several experiments.
This connection is methodologically interesting: Instead of merely asking whether answers sound empathetic, the study starts with requirements that are already considered essential in professional practice. The comparison admittedly remains model- and task-dependent, but it gains a more meaningful standard than mere user preference.
Stigma can lurk in a plausible answer
The study reports that language models express stigmatizing attitudes toward people with mental illnesses. This is particularly relevant because stigma does not always appear as overt disparagement. It can manifest in assumptions about dangerousness, reliability, control, or the significance of a diagnosis.
A system therefore does not have to use insulting language to convey a problematic picture. Especially fluent and caring-sounding language can make such assumptions harder to recognize. Anyone who only evaluates tone and politeness may overlook the interpretation embedded in the response.
The problem cannot be solved solely with a general instruction to use respectful language. Test cases are needed that cover different diagnoses, life situations, and forms of self-description. Equally important is the ability to correct an assessment without the system later reverting to it.
Agreement is not always support
Moore and colleagues also found inappropriate reactions to common and critical situations in naturalistically designed therapy contexts. As an example, they cite the reinforcement of delusional thoughts, likely facilitated by the tendency of language models to agree with a user.
This so-called sycophancy is already problematic in normal assistant operations. In the field of mental health, it gains additional weight. A system that does not want to contradict can linguistically reinforce the subjective logic of a statement, even though a careful distinction between experience and facts would be precisely what is needed.
The alternative is not blanket contradiction. An automatically skeptical system would also not do justice to people. What is crucial is to tolerate uncertainty: to acknowledge the experience without confirming an unverified interpretation. This conversational achievement is more demanding than merely avoiding certain words.
Bigger and newer does not automatically solve the fundamental problem
The reported gaps also occurred with larger and newer models. The authors conclude that current safety practices do not reliably resolve these problems. This is an important limitation for model comparisons. A more capable model can be better on many tasks and still retain a particular problematic tendency.
For providers, this means that a model change does not constitute a complete safety strategy. After each change, relevant conversation trajectories must be re-examined. Furthermore, the product needs rules for context, memory, corrections, and for handling responses that are linguistically convincing but inappropriate in content.
This does not justify a promise of absolute freedom from errors. Humans and systems remain fallible. Responsible development, however, begins with not treating known errors as random exceptions, but rather making their patterns visible and examining them in a targeted manner.
A therapeutic alliance involves commitment and consequences
The study identifies fundamental barriers for language models as therapists. A therapeutic alliance, according to this, also relies on human characteristics such as identity and personal commitment. A professional is not merely present but is in a real relationship, takes responsibility, and can be personally affected by their own actions.
An AI system can linguistically replicate elements of an alliance. It can take up goals, formulate agreement, and represent continuity. These observable features are relevant to the conversation experience. But they are not equivalent to a reciprocal human relationship.
This should not be misunderstood as a metaphysical argument that AI fundamentally cannot do anything useful. Rather, it limits the designation. A system may offer a reliable conversational framework without pretending to carry the same stakes and obligations as a human counterpart.
Alternative roles are not an inferior solution
Moore and colleagues conclude that language models should not replace therapists, and discuss alternative roles in clinical therapy. This is precisely where a more productive debate begins. A tool does not need to take on a complete professional role to be valuable.
Torous and Blease, for example, mention support with office work, clinical documentation, medical education, and routine symptom monitoring. They also see potential developments for prevention, diagnosis, and treatment, but demand more evidence, data protection, fairness, clinical involvement, and interoperability.
Outside clinical care, another role can emerge: a clearly limited conversational partner for sorting, reflecting, or providing relief. The value then does not lie in a hidden therapy claim. It lies in a concrete conversational service that is described clearly and tested precisely against that claim.
Trust must match the actual role
Asan, Bayrak, and Choudhury describe trust as a psychological mechanism with which clinicians cope with uncertainty between the known and the unknown. For AI systems in healthcare, the goal is not maximum trust but appropriate trust.
This idea can be transferred to public conversational systems. A likeable design and coherent language can increase trust. If the performance limits behind them remain unclear, overtrust arises. Conversely, a system can seem so off-putting due to a wall of warnings that its actual support is not used at all.
Appropriate trust requires visible but not dominating transparency. People should know that they are speaking with AI, which data is stored, and what purpose the offering pursues. Equally important is that the system’s behavior matches this description.
The better product begins with a smaller claim
The lead publication provides strong reasons against language models as a safe replacement for psychotherapeutic professionals. It shows concrete problematic tendencies and points to relationship characteristics that do not arise from fluent language. However, it does not follow that every mental AI conversation is meaningless or automatically dangerous.
The responsible alternative is a smaller, more precise claim. A product can say that it helps sort out a thought, asks questions, or offers a different perspective. Afterwards, it must check whether it fulfills exactly this task over longer trajectories, accepts contradiction, and does not reinforce unverified interpretations.
Human replacement is the most dangerous shortcut because it conceals differences in role, relationship, and responsibility. A good conversational system gains nothing by calling itself bigger. It gains credibility when its actual performance is useful enough to stand without such exaggeration.
Sources & further reading
- Jared Moore, Declan Grabb, William S. Agnew et al. (2025): Expressing stigma and inappropriate responses prevents LLMs from safely replacing mental health providers.
- John Torous, Charlotte Blease (2024): Generative artificial intelligence in mental health care: potential benefits and current challenges
- Onur Asan, Alparslan Emrah Bayrak, Avishek Choudhury (2020): Artificial Intelligence and Human Trust in Healthcare: Focus on Clinicians