Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
The system used existing peer responses
Instead of generating free empathetic texts, the agent drew on an existing pool of responses from online peer support. As a result, the linguistic material originally came from real human support.
Information retrieval methods and word embeddings searched for responses that were intended to match the expressed concerns in terms of content. The approach combines human-generated content with automatic selection.
This construction reduces certain risks of free generation, such as newly invented details. It creates other problems: a response can be semantically similar and yet inappropriate in the specific biographical context.
Rights, consent, and purpose limitation of historical peer data are also central questions for today's applications. Human texts are not automatically freely usable training material.
79.2 percent of the responses were considered acceptable
Of 3,770 agent responses, users rated 2,986 as acceptable. The large total of 37,169 ratings gives the result a broad perceptual basis.
Acceptable, however, is a lower standard than particularly helpful, precise, or equivalent to human support. Roughly one fifth of the responses did not even reach this threshold.
For emotional conversations, this remainder is relevant. An inappropriate response can have a stronger impact after a vulnerable disclosure than an error in an ordinary service inquiry.
Evaluation should therefore not only report the average but also examine the types and consequences of unacceptable responses. Wrong tone, loss of context, and problematic interpretation require different improvements.
In addition, the distribution across individuals matters. If some users almost always receive fitting responses while others repeatedly receive inappropriate ones, a good average can conceal inequality. Language, topic, and style of expression should therefore be examined as possible influencing factors.
The source changed perceptions of identical texts
Users significantly preferred their peers’ responses over the agent’s efforts. In the controlled experiment, the effect persisted even though the response itself was identical.
The mere labeling as human or machine-generated influenced the judgment. People interpret empathy not only from words, but also from the assumed experience and intention of the counterpart.
A human response can be understood as testimony of lived experience. A system can reproduce the same sentence, but it does not possess a comparable biography or emotional involvement.
Clear AI labeling can make this perceptual disadvantage visible, but it is ethically indispensable. The goal must not be to hide the machine origin in order to obtain better evaluations.
Emotional generation was further developed technically
Zhou, Huang, and Zhang introduced the Emotional Chatting Machine, an approach intended to generate not only relevant and grammatical content but also emotional consistency.
The system modeled emotion categories, a changing internal emotional state, and an external emotion vocabulary. Experiments showed responses that could be appropriate in both content and emotion.
This technical representation is not experienced emotion. It controls linguistic patterns that people can perceive as emotionally appropriate.
For product communication, this distinction is important. A model can control emotional expression without feeling. The benefit lies in appropriate communication, not in an asserted inner state.
STEF uses the conversation trajectory rather than just a single turn
Wang, Peng, and Zha criticize that emotional support is often derived from only a single current message. As a result, subtle changes across the conversation trajectory are lost.
Their STEF agent combines an emotional fusion mechanism with an encoder for the development of support strategies. It is designed to consider earlier emotions and the strategy trajectory jointly.
On the ESConv benchmark, STEF performed better than competitive baselines. This demonstrates the technical value of trajectory signals in strategy selection.
A benchmark result does not yet provide evidence of clinical efficacy. It shows that systematically considering past turns can improve the prediction of supportive responses.
The trajectory is more than an emotion curve
In real conversations, mood is not the only thing that changes. People clarify facts, reject suggestions, switch goals, or explicitly state how the chat should respond.
A system that only stores an emotional tendency can miss a central boundary in the conversation. “I want to gather first” is not an emotion but an instruction for how to proceed in the interaction.
Good state management should therefore map emotion, conversational intent, corrections, and open content separately. Each level has a different validity period.
The visible proof is simple: does the new information actually change the next response? An extensive internal state is worthless if the system continues the same routine.
Human content does not automatically resolve responsibility
Using real peer responses can introduce natural language and patterns of experience. However, it can also reproduce inappropriate advice, cultural assumptions, or historical errors.
Selection models therefore require content review, contextual boundaries, and a way to remove problematic templates. A suitable semantic similarity is not sufficient.
Generative models can flexibly adapt a template but increase the risk of new fabrications. Hybrid systems must document which part originates from the source, the selection, and the generation.
For sensitive support, the user should not believe that a current other person has read their message when in fact only a historical text was automatically selected.
Even a technically perfect source attribution does not answer whether the original authors agreed to this further use. Corpus governance is therefore part of system quality and not merely a legal footnote.
Empathy remains a relationship between text and expectation
The peer-support study provides a rare large dataset of user judgments and controlled evidence that the claimed source changes the evaluation of identical responses.
ECM and STEF show how emotional categories and trajectory can technically be incorporated into response generation. These methods improve expression and strategy, not the ontological question of whether a machine feels.
A responsible chatbot should therefore not claim human empathy. It can offer attentive, well-reviewed, and context-aware responses and examine their impact with users.
79 percent acceptable is a considerable start, not a final grade. What remains crucial is which responses fail, how the system learns from corrections, and whether its clearly named technical role still becomes helpful for people.
Sources & further reading
- Robert R. Morris, Kareem Kouddous, Rohan Kshirsagar (2018): Towards an Artificially Empathic Conversational Agent for Mental Health Applications: System Design and User Perceptions
- Hao Zhou, Minlie Huang, Tianyang Zhang (2018): Emotional Chatting Machine: Emotional Conversation Generation with Internal and External Memory
- Qing Wang, Shuyuan Peng, Zhiyuan Zha (2023): Enhancing the conversational agent with an emotional support system for mental health digital therapeutics