Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

An experiment involving personal stress

Lee and Hahn asked participants to write about stressful interpersonal events. The chatbot asked questions about these experiences and, depending on the experimental condition, provided either informational or emotional social support.

Informational support consisted, in principle, of useful information and advice. Emotional support worked through empathy and encouragement. In addition, the researchers examined the extent to which participants explicitly or implicitly attributed a human-like mind to the chatbot.

In this way, the study did not merely test which response was better liked. It examined the interplay between the type of support and the notion of what kind of counterpart a chatbot actually is.

Explicit anthropomorphization increased helpfulness

When participants explicitly attributed a human-like mind to the chatbot, they rated its support as more helpful for coping with the distressing event. The social interpretation of the system thus influenced the effect of its message.

This is easy to understand. In everyday human life, empathy presupposes a counterpart that can at least in principle understand, sympathize, or care. If the same wording is perceived as statistically generated text, it can appear empty or contrived.

A product cannot simply derive from this that it should anthropomorphize as strongly as possible. A higher perceived efficacy is not automatically informed agreement. The stronger the staging of a feeling counterpart, the more important transparency and protection against false expectations become.

Emotional support could even be harmful

The most striking finding concerns implicit attribution. When participants did not implicitly attribute a human-like mind to the chatbot, emotional support reduced the effectiveness of the message. This effect did not occur with informational support.

"Harm" here does not mean proven clinical harm. It means that the message was rated as less effective for the stressful event. Nevertheless, the result is significant for conversation design: more empathetic language can make a response worse.

The problem likely lies not only in the content but in the mismatch between source and tone. When a machine speaks as if it were inwardly empathizing, this can seem implausible to some people and overshadow the actual help.

Information does not require a feeling sender

Useful information can help regardless of whether it comes from a human or a system. A hint, a comprehensible explanation, or an organized selection is judged by its content. This makes informational support more robust against perceptions of the chatbot.

This robustness is attractive for mental-health conversational systems, but it must not be mistaken for general superiority. Someone who wants to talk things out does not immediately need advice. A factually correct list can be just as inappropriate at the wrong moment as an artificially intimate consoling sentence.

The lesson, therefore, is not “always inform.” It is: The type of support must match the concern, and its linguistic form must match the recognizable nature of the system.

Information can also be very small in this context. A clear answer to a specific question does not have to grow into an action plan. Especially in personal conversations, quality is shown by the chat fulfilling the requested function and then being able to listen again, rather than turning every useful hint into a counseling sequence.

Everyday users nevertheless report emotional help

Ta and colleagues found diverse forms of social support in 1,854 public Replika reviews and open-ended responses from 66 users. People described companionship, a non-judgmental space, encouraging messages, and positive feelings.

This qualitative finding does not contradict the experiment. Replika users are a self-selected group with concrete experiences and possibly a different relationship to the system. The experiment shows a conditional effect; the experience reports show that emotional support is indeed accepted under certain circumstances.

Together, both works make the differences between people visible. What is comforting for one person feels artificial to another. A uniform empathy style cannot capture this diversity.

Attachment changes credibility

Xie and Pentina analyzed interviews with 14 Replika users. Under stress and in the absence of human companionship, attachment could develop when responses were experienced as emotionally supportive, encouraging, and psychologically safe.

In an established usage history, an emotional response has a different context than in the first contact. Earlier fitting reactions, personalization, and routine can increase credibility. Conversely, a sudden standard comfort phrase after many individualized conversations can be particularly disappointing.

Empathy is thus not an isolated property of the text. It emerges from the course of the conversation, expectations, and fit. A sentence can look linguistically empathetic and still be inattentive in the specific dialogue.

Naturalness does not require claims of feeling

A chatbot does not have to become formal or cold in order to remain honest. It can respond in everyday language, pick up on a thought, and create closeness through attention rather than through claimed feelings. “Tell me more” does not require a pretense of one’s own experience.

Particularly delicate are formulations such as “I feel your pain” or “I am worried about you.” They attribute inner states to the system that it does not possess. A more restrained language can fulfill the same conversational function: taking what was said seriously without claiming a human inner world.

Clear AI labeling can also be combined with a good tone in this way. Transparency and naturalness are not opposites, as long as the chat does not constantly push technical explanations into the conversation.

The Right Support Outperforms Maximum Empathy

The study by Lee and Hahn reveals a concrete limitation of the widespread empathy ideal. Emotional language does not work independently of what people attribute to the sender. In the absence of implicit anthropomorphism, it can even weaken the message.

Good conversational design should therefore not select the most emotional responses, but the most fitting ones. Sometimes that is a brief reflection, sometimes factual information, sometimes a follow-up question, and sometimes simply space for the next sentence.

Artificial empathy fails when it is imposed on the conversation as a role. Credible support emerges when the system listens carefully, does not overdo its manner, and lets the user decide what form of help is actually wanted at any given moment.

Sources & further reading