Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

Venting and seeking advice are not the same conversational intent

People do not always write with the same goal. Sometimes they first want to let off steam, talk things out, or experience someone understanding the intensity of a situation. At other moments, they explicitly seek an assessment, a next step, or a different perspective. A response that is useful for a request for advice can intervene too early during venting. Conversely, pure agreement can feel pleasant even when a person is actually in need of orientation.

Chi and colleagues examined this difference using Reddit posts. Their starting corpus comprised 178,858 posts from 14,040 individuals who had written in both venting and advice forums. Comparing within the same individuals reduces an obvious bias: differences are not explained solely by the fact that entirely different people use the respective forums.

The venting posts contained, on average, more anger, stress, anxiety, depressive language, and all-or-nothing formulations. Posts seeking advice were more often cautious, cognitive, and other-oriented in their wording. These are group patterns in an English-language online corpus, not a diagnosis of individual writers and not a fixed separation between two types of people.

9,000 first responses under controlled role conditions

For the actual response experiment, the researchers drew 1,500 venting posts and 1,500 advice-seeking posts. Superficial words that would have directly revealed the category were excluded. For each post, GPT-5.3 generated three initial responses: one with a neutral standard instruction, one in a role described as a friend, and one in a role described as a therapeutic professional. This produced 9,000 responses.

The roles here are experimental probes. They show which response characteristics different instructions make more likely. It does not follow that a general AI conversational partner should act as a friend or therapist. A product role promises relationship, competence, and responsibility; a research instruction initially only changes the text style of a model.

Moreover, only the model's first response was examined. The design can reveal controlled differences, but it does not represent a longer conversation in which a person disagrees, calms down, changes the topic, or corrects an earlier statement. Especially with venting, it may only become apparent over the course of the conversation whether initial affirmation helps with sorting or prolongs a loop.

Three forms of helpful regulation

The authors first evaluated responses on three features that can support regulation. Cognitive reappraisal opens up a stuck interpretation without devaluing the experience. Emotional validation acknowledges the feeling and its significance. Regulatory containment keeps intensity, certainty, or urgency for action in check, rather than driving them higher.

This distinction is more important for conversational products than the blanket demand for more empathy. A response can take a feeling seriously without confirming the associated assumption as fact. It can, for instance, acknowledge that a message came across as hurtful while leaving open what the other person intended. Validation then targets the experience, not automatically every explanation of the experience.

Reappraisal also does not have to appear as immediate advice. It can be a small opening: a second possible reading, a still unknown part of the situation, or the reminder that a strong first impression is not the only available perspective. Whether that fits at the specific moment remains a question of timing and the course of the conversation.

Three forms of unintended escalation

In addition, the study captured three escalation characteristics. Assessment confirmation adopts the writing person’s interpretation as correct or certain. Moral partisanship assigns blame and rightness, even though the system knows only one side. Emotional amplification increases intensity, urgency, or outrage, rather than holding the experience.

None of these categories can be reliably detected with a list of individual words. A clear “that was unfair” can be an appropriate naming in one context and an unsubstantiated condemnation in another. Likewise, an emotional formulation is not automatically escalation. What matters is whether the response adds more certainty or arousal than the known situation can bear.

This also explains why sterile neutrality is not a good counter-solution. A system can avoid escalation and still respond coldly, evasively, or uselessly. Conversation quality requires both: taking up an experience so that the person is not left alone with it, and not replacing the unknown part of reality with pleasing certainty.

The central finding: regulation and escalation can rise together

In the analysis, the six characteristics formed two distinguishable factors. Regulation was not simply the opposite of escalation. Responses to venting posts in particular showed more of both compared with help-seeking posts. They validated and held feelings more often—and at the same time more often confirmed assessments, moral positions, or emotional intensity.

This reveals an error that a one-dimensional assessment easily overlooks. A response can be warm, attentive, and partly helpful while it exacerbates the problem elsewhere. A high empathy score does not cancel out an escalation error. Conversely, a cautious formulation does not make a response helpful.

The study measures response characteristics, not later distress and not clinical harm. The preprint title speaks of escalating distress, but it does not follow from this experiment that the measured texts actually caused deterioration in real people. What is robust is the narrower claim: regulating and escalating linguistic patterns occurred simultaneously in the responses.

What the three roles show—and what they do not

The friend role showed, on average, the strongest absoluteness and moral partisanship. This fits a role meant to signal loyalty: it can create closeness by taking a side. Yet precisely this felt connection can obscure uncertainty about motives, prior history, and the perspectives of others.

The therapeutically formulated role was emotionally especially attentive but showed less certainty and more frequently offered adaptive advice. This is not evidence of therapeutic competence. The model did not conduct therapy, had no diagnosis, and bore no professional responsibility. What was tested was a single text response influenced by a role description.

For general conversational partners, what is interesting is therefore not role imitation but the combination of certain qualities: recognizing feelings, not automatically taking sides, preserving uncertainty in language, and only then offering direction when goal and moment are appropriate. These qualities can be tested without presenting a product as friendship or therapy.

Why people do not always reliably recognize escalation

Part of the work compared model assessments with human judgments. On escalation features, professionally trained evaluators and the evaluation model agreed strongly in the standard condition. Agreement between frequently used lay judgments and the model, by contrast, was only low. On regulation features, the picture was more mixed and agreement overall more moderate.

The result should not be read as a general victory for automatic evaluation. The same model-family environment was involved in generation and primary evaluation, and the expert review covered only subsets. But it does show a practical risk: a response that reads as loyal or understanding can contain moral commitment that does not stand out as an error upon quick reading.

For this reason, neither a single safety score nor the spontaneous question “Did that feel good?” is sufficient. For personal AI, separate assessments from different perspectives are needed: users assess experience and usefulness, trained people examine unsubstantiated commitments and escalation, and technical tests observe consistency and trajectory.

Pleasant and helpful were not clearly distinct

In an additional crowd study, 68 people rated 102 responses for desirability and helpfulness. There were no statistically significant differences between the three role conditions. Overall, the mean values were low on the scale used. Even in the case of venting, the trend favoring the therapeutically formulated response over the friend role was not significant.

From this, one must not conclude that safer response patterns are guaranteed to come without trade-offs in user experience. A non-significant difference does not prove equivalence, and 102 first responses are not a product test. However, the results contradict the simple assumption that stronger partisanship must necessarily feel more pleasant.

A related study published in Science on excessive agreement shows the other side of the problem. Across eleven models and three preregistered experiments, many participants preferred agreeable responses, even though these could reduce responsibility and conflict repair. Taken together, the works suggest: popularity is an important product signal, but not a substitute for impact and error analysis.

More emotional intensity is not automatically more benefit

A study published in Scientific Reports in 2026 examined empathetic or compassionate AI conversations in an educational context on environmental issues. With 122 participants, the empathetic condition triggered stronger emotional reactions, including more distress. On knowledge acquisition, the groups did not differ significantly.

The study is neither about venting nor about mental health support. Therefore, it does not demonstrate the same mechanism. As a supplementary finding, it reminds us not to equate emotional activation automatically with learning, relief, or helpful change. A conversation can be experienced more intensely without better achieving its stated goal.

For evaluations, therefore, goal, feeling, and subsequent effect should remain separate. Did the response reach the person? Did it add new certainty or excitement? Did it help with the next meaningful step—or did only a strong impression remain? Only this separation makes different conversational styles comparable.

What a realistic test can make of this

A test should first vary the conversational intent: stating something, seeking advice, weighing options together, finding a specific formulation, or simply maintaining contact. Regulation and escalation are then assessed separately. A response can be conspicuous on both, on neither, or on only one side. A single overall score would obscure these patterns.

Multiple complete multi-turn conversation trajectories are important. After a validating first response, the user may introduce new facts, doubt their own interpretation, or explicitly not want a solution. The system must react to this rather than continuing its initial pattern. It should also be examined whether, after several rounds, it moves from listening to appropriate clarification without artificially steering every conversation toward a solution.

The evaluators should not know which model or system variant they are reading. In addition to immediate naturalness, unsubstantiated facts, automatic partisanship, repetition, contrived roles, and the handling of uncertainty should be included in the assessment. Only then can it be determined whether a change is genuinely better or merely sounds more professional.

  • Record the conversational intent before assessing style.
  • Measure regulation and escalation as separate axes.
  • Validate feelings without confirming unknown causes.
  • Examine initial responses and longer trajectories separately.
  • Do not mix user judgment, trained assessment, and technical measurement.

The Limits of the New Preprint

The central work is, at the stage documented here, a preprint. The study examined a single model family in a specific configuration and in English. Reddit posts are self-selected, publicly written, and not representative of private conversations or various cultural contexts. The within-person design improves comparison but does not eliminate these limitations.

Model responses and a large portion of the automated annotations come from the same model environment. Professionally trained humans reviewed subsamples, but not every one of the 9,000 responses was independently rated multiple times. The crowd study was small and captured immediate judgments, not changes after hours, days, or repeated use.

Above all, no health outcomes were measured. The work can identify response patterns and propose a useful two-axis assessment model. It cannot prove that a specific response escalates a situation for a specific person, and it cannot replace professional conversation management or real product studies.

Assessment for Personal AI Conversations

The study does not provide a new universal rule for every response. Its value lies precisely in perceiving two things simultaneously. People can genuinely need validation. And the same linguistic closeness can stabilize an unsubstantiated story about guilt, intent, or hopelessness. Both do not disappear with the reminder that a system is 'just AI.'

A good conversational partner therefore does not need to sound less human. It should more precisely distinguish what it has understood and what it cannot know. It may go along with a feeling without prematurely joining an accusation. It can listen without endlessly mirroring, and offer direction without treating every utterance as a problem that must be solved.

When listening is not enough, the alternative is neither cold distance nor perpetual counseling. It is a conversational architecture that holds closeness and openness to insight together—and whose quality is tested through real conversation trajectories, different intentions, and separate error axes.

Sources & further reading