Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

155 conversations reveal an apparent paradox

The study examined the effect of chatbot identity and perceived empathy on the overall conversation experience. The basis was 155 conversations from two existing datasets.

GPT-based chatbots were rated significantly higher in conversation quality. At the same time, they were considered less empathetic compared to human conversation partners. Good linguistic quality and experienced empathy thus did not coincide.

GPT-4o was additionally used for annotations. Its empathy judgments agreed with user ratings in rating chatbots lower than humans.

The number of conversations and the datasets used limit generalizability. Nevertheless, the separation is conceptually important: readability, coherence, and responsiveness do not fully capture whether someone feels seen over the course of the conversation.

Empathy is attributed, not built in

A system can produce a formulation intended to be empathetic. Whether it comes across as empathetic to the other person depends on the situation. The same sentence can seem comforting, routine, or inappropriate.

When a user gives a factual correction and the chat again interprets things emotionally, even a warm response comes across as inattentive. When someone simply wants to tell their story and receives a solution after every sentence, friendly language cannot mask the breakdown in the conversation.

Empathy is therefore not a fixed property of text. It arises in the relationship between utterance, expectation, and reaction. For technical systems, it is a perception on the part of users, not an internal state.

This distinction also guards against misleading product claims. A chat can be developed and tested with an eye to empathetic communication. It should not claim to experience feelings in a human way.

Earlier experiments nonetheless show the value of emotional language

In two experiments, Liu and Sundar compared sympathy, cognitive empathy, and affective empathy with purely factual health information. 158 people read a dialogue; another 88 interacted with a real chatbot.

Expressions of sympathy and empathy were preferred over emotionless advice. This was especially true for people who were initially skeptical about whether machines possess social cognitive abilities.

This does not contradict the more recent study. Emotional signals can improve a conversation without reaching the level of humanly experienced empathy. The appropriate conclusion is therefore neither “empathy formulas do not work” nor “more empathy formulas solve the problem.”

Dosage is more important. A brief acknowledgment can be helpful on a sensitive topic. A fully formulated emotional monologue by the system, on the other hand, can shift attention away from the user and toward the machine’s purported reaction.

Conversation quality has multiple levels

A good conversation can be understandable, relevant, respectful, and coherent. It can recall earlier statements, maintain a chosen tone, and handle uncertainty appropriately. Empathy is an important dimension, but not the only one.

A system can sound empathetic while giving factually incorrect advice. Conversely, a factually correct answer can be emotionally cold. Evaluation should assess such dimensions separately so that a high overall score does not mask a specific weakness.

Conversation function also matters. For an informational question, clarity is paramount. When the goal is relief, too much explanation can be disruptive. When sorting things out together, a cautious follow-up question may be more helpful than immediate affirmation.

Chat offerings therefore need more than a universal empathy prompt. They require rules for conversational direction, course of the conversation, and transitions between listening, asking, interpreting, and suggesting.

Individualization must not merely be claimed

In the rhinoplasty study by Xie and colleagues, ChatGPT answered nine questions from a professional checklist. Specialists rated the responses as coherent, accessible, and informative.

The responses emphasized the importance of an individualized approach, but themselves offered only limited detail and personalization. The example illustrates a common difference between a correct general statement and actual adaptation.

Similarly, a chatbot can write that everyone experiences a situation differently, and in the next sentence still use a standard template. The claim of individuality is not evidence of individualized behavior.

Adaptation only becomes measurable through variations: Does the system respond differently when the goal, boundary, or context changes? Does it carry forward a correction? Does it avoid the same recommendation after it has been rejected?

Too much warmth can create new risks

Emotionally intense language can intensify a closeness that the system cannot sustain. Particularly problematic are exclusivity, messages of dependency, or the portrayal that only the chatbot understands the person.

A credible conversational partner may be friendly, casual, or calm, without feigning human reciprocity. The AI labeling should remain clear and not disappear behind a heavily anthropomorphized character.

Agreement is also not the same as empathy. Sometimes respectful disagreement is more helpful than reflexive affirmation. What matters is whether the chatbot understands the perspective without adopting every conclusion.

A good boundary does not have to sound cold. It can be phrased naturally while still preventing the system from promising a relationship, memory, or responsibility that does not exist technically or organizationally.

Trajectory tests are more informative than sample sentences

A single screenshot shows that a system can produce a fitting response. It does not show whether the system remains attentive over ten or twenty turns. That is precisely where helper drift, repetitions, and forgotten corrections arise.

Test cases should therefore contain small changes: the user qualifies a statement, rejects a suggestion, wants to keep talking, or shifts the direction of the conversation. The quality of the response must visibly adapt to these changes.

In addition to expert assessments, user judgments belong in the evaluation. People are best able to report whether a reaction felt helpful, intrusive, or inauthentic. Experts additionally examine boundaries, risks, and conversational techniques.

Automated evaluations can presort large volumes, but they should not be the sole arbiter of what counts as empathy. The study itself at least shows that a model rating can align with user judgments; this alignment must be re-examined for each new task.

The goal is fit rather than maximal empathy language

The analyzed conversations show that modern chatbots can achieve high overall conversation quality and yet be perceived as less empathetic than humans. This is not merely a weakness of individual words, but a property of the entire interaction.

Product teams should therefore not optimize for the number of emotional formulations. They should check whether the chat listens, tolerates uncertainty, respects boundaries, and adapts its response to the desired function.

Sometimes the best response is just a brief sentence and room for the next message. Sometimes a follow-up question is necessary. And sometimes the user wants a clear opinion rather than further mirroring.

Empathetic sentences can be a helpful resource. Empathetic experience only emerges when tone, content, and course of the conversation fit together. It is precisely this fit that a conversational system must visibly learn and repeatedly prove under realistic conditions.

Sources & further reading