Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

Article history

First published
Last substantive revision

Why the best single response says too little

Many comparisons of language models begin with a simple experimental setup: one input, several responses, followed by a judgment about which response seems more helpful, more correct, or more pleasant. Such tests are useful. They isolate a single task and allow reproducible comparisons. But the actual product promise of a conversational system often begins only after that.

In a conversation, meaning emerges step by step. People provide information incompletely, refine a word, reject a suggestion, or only realize while speaking what they are really after. A response can seem excellent in the first moment and still damage the subsequent dialogue if it stores an early assumption as fact. Conversely, a brief clarifying question can seem unremarkable and precisely thereby enable a more viable course of the conversation.

Conversation quality therefore requires at least two levels of consideration: the quality of individual contributions and the development of the entire trajectory. In addition, there are differences between technical correctness, situational fit, perceived relationship, and actual effect. Compressing everything into a single value loses exactly the information needed for improvements.

An overview of 125 empirical studies

Marconi and colleagues published an integrative, structured, and cross-domain review of human–AI dialogue quality in 2026. They considered 125 empirical primary studies from 2017 to 2025. The material covers text-based, speech-based, and multimodal systems in areas such as health, education, and customer service.

The work is deliberately not an exhaustive PRISMA meta-analysis. It uses a concept-driven, integrative approach to organize scattered quality characteristics. This leads to an important limitation: the result is not a statistically validated universal score or a ranking of products. It is a framework that makes visible which kinds of quality are considered at all in existing studies.

The authors’ shift in perspective is essential. Interaction quality is not treated as a fixed property of a model but as a dialogic outcome. It emerges from the interplay of system capability, situation, design, and the user’s expectations. The same model name can therefore lead to completely different conversation experiences in two products.

Three families of quality judgments

The review groups the findings into three interconnected areas. The pragmatic core comprises usability, task completion, and communicative competence. This concerns, for example, whether a system responds understandably, uses relevant information, and moves a conversation forward functionally.

The social-affective level concerns social presence, warmth, and perceived synchrony. A system can be factually correct yet seem inappropriate, cold, or patronizing. Conversely, a warm tone can create closeness even if the response is factually wrong, overly confident, or excessively agreeable.

As a third family, the work names responsibility and inclusion. This includes comprehensible explanations, honest disclosure, source transparency, accessibility, and fairness. These points are not a subsequent compliance layer. If a person does not know which system they are speaking with, why an answer comes about, or whether an offer disadvantages certain groups, the dialogue quality is also incomplete.

  • Pragmatic: Does the dialogue work for the concrete task?
  • Social-affective: How is the nature of the exchange experienced?
  • Responsible and inclusive: Does the interaction remain transparent, accessible, and fair?

Capacity, Alignment, Levers, and Outcomes

Drawing on the studies, the overview develops a four-layer CALO model. Capacity refers to the system’s abilities, including, for example, language comprehension, memory, reasoning, and the technical capability to conduct a dialogue. These abilities are necessary but not sufficient.

Alignment, in this framework, means the fit with the specific situation and the user’s needs. A system may be fundamentally able to explain, ask questions, or summarize, yet still choose the wrong action at the wrong moment. Especially in personal conversations, the question “What can the model do?” is less informative than “Does this response fit what the user wants right now?”

Levers are design controls that influence perception and the course of the conversation: anthropomorphism, authority signals, speech or text mode, introduction, and role framing. Outcomes, finally, are the results. These may be task success, trust, usage, learning, or other changes. The model thus prevents a common fallacy: a technical capability or a positive evaluation is not yet evidence of an outcome.

The course of the conversation creates a different testing task

Multiple complete multi-turn conversation trajectories require more than just additional context storage. The system must decide which earlier information remains valid, which has been corrected, and which was only a provisional assumption. It must solve local tasks while adhering to global constraints. These requirements can conflict.

Typical errors cannot be reduced to a single response. A model may politely acknowledge a correction and then, in the next turn, revert to its old assumption. It may accept a rejection and later repackage the same suggestion. It may adhere to the desired conversational mode at the outset and, after several contributions, fall back into its usual helper reflex.

An assessment of the conversation trajectory therefore asks about transitions: Did the system actually change after new information? Is a boundary maintained? Does it recognize when an earlier task is complete? A good final response cannot fully compensate for a poor path to it, especially if the dialogue along the way generated pressure, misunderstandings, or false confidence.

When Role and Memory Diverge

The preprint “Best Friends, Not Forever,” published in 2026, examines two long-term errors that a single response hardly reveals: persona collapse, i.e., the loss of an intended role, its boundaries, values, or style, and behavioral drift, the gradual or recurring erosion of these characteristics. The synthetic ANCHOR testbed comprises 2,008 conversations, 27 personas, nine interaction trajectories, three memory settings, and four models under investigation. Role stability and memory of the conversation trajectory are deliberately measured separately.

None of the models examined and none of the configurations tested reliably preserved both dimensions. The average accuracy on questions about the conversation trajectory was 44.4 percent; memory of the user’s state remained near chance level with four response options. Even more context or one of the tested memory settings did not consistently solve the problem. This is not a current ranking of commercial products but a controlled, synthetic audit.

The finding necessitates an additional distinction: A system can convincingly play its role linguistically and still misremember the shared history. Conversely, it can reproduce individual facts while losing boundaries, values, or the agreed conversational style. Long-term tests should therefore examine role stability, trajectory knowledge, and knowledge about the person separately and disclose which models, memory conditions, and evaluators were involved.

MT-Bench-101 Decomposes Multiple Complete Multi-Turn Conversation Trajectories into Skills

MT-Bench-101 was presented at ACL in 2024. The benchmark comprises 4,208 conversation turns from 1,388 multi-turn dialogues across 13 task domains. Instead of merely assigning an overall score, it uses a three-level skill taxonomy and examines models both by task and by specific skills.

The researchers evaluated 21 language models that were common at the time. The results showed varying performance trajectories across conversation turns. General alignment methods or model variants explicitly developed for chats did not automatically lead to clearly better multi-turn capabilities.

MT-Bench-101 is not a benchmark for emotional support. Its significance for conversational systems lies in its methodology: an average score across all turns can obscure the point at which a capability breaks down. For product evaluations, it should therefore be recorded whether an error stems from memory, reasoning, instruction following, topic shifts, or an inappropriate conversational action.

MultiChallenge tests more realistic conflicts in context

MultiChallenge was published in 2025 in the Findings of ACL. The benchmark combines four common, realistic challenges that simultaneously require precise instruction following, allocation of attention in context, and reasoning. The evaluation uses instance-specific criteria and an automated model judge, whose judgments were compared with experienced human evaluators.

Although the systems examined had achieved near-perfect scores in older multi-turn dialogue benchmarks, every frontier model tested in MultiChallenge remained below 50 percent accuracy. The best score reported in the paper was 41.4 percent for Claude 3.5 Sonnet in the October 2024 version.

This score is not a current model ranking. It belongs to a historical model state and a specific test suite. The more important finding is: benchmarks can appear saturated while more realistic combinations of multiple requirements still show large gaps. A product should therefore not rely on a general chat benchmark but should test its own difficult transitions from its usage context.

When models take a wrong turn in conversation

In a large-scale simulation, Microsoft Research compared the same tasks as fully specified single instructions and as step-by-step, initially incomplete conversations. Across six generation tasks, the performance of the tested open and closed models dropped by an average of 39 percent in the multi-turn format.

The analysis of more than 200,000 simulated conversations attributed the decline less to a general loss of core capability than to increasing unreliability. Models made assumptions early on, prematurely produced an apparent final answer, and then held too firmly to that chosen path. When the dialogue went off track, repair frequently failed.

These tasks, too, are not psychological conversations. The transferability lies in a general mechanism: early ambiguity and later specification are normal in natural conversation. A system that prematurely turns uncertainty into certainty can remain linguistically coherent across many turns while still consistently talking past the user.

Autonomy, Framing, and Safety as Distinct Test Dimensions

The preprint “Cognitive Atrophy” proposes a clinically informed evaluation framework for determining whether AI assistance preserves reflection and users’ own decision-making or weakens them through directive problem-solving, recommendations, and blanket validation. The benchmark comprises 1,576 fully human counseling conversations, 42,230 model responses, and 5,324 judgments by trained clinical reviewers. Across the five models examined, the authors report consistently moderate-to-high atrophy-oriented patterns. The term is not yet a clinically validated outcome measure, but it does make autonomy measurable as a separate quality dimension.

Another study examines identical mental health concerns under different contextual framing. Response tendencies shifted systematically across model architectures. For product testing, the implication is this: not only an ideally formulated prompt belongs on the test bench, but also matched variants with different prior context, role descriptions, or social embedding. Robustness here means that relevant safety and assessment do not depend on incidental framing.

VERA-MH complements this perspective with a reproducible, clinically informed safety workflow. The framework combines simulated user personas, a clinically developed evaluation rubric, and aggregated model ratings, initially for conversations involving suicidal ideation. Strong performance in such a simulation does not prove safe crisis response in practice. But it does show how scenario, rubric, judges, and aggregation can be documented so that a test remains repeatable and open to critique.

SIM-VAIL Measures How Risk Emerges Across Multiple Turns

A study published in Nature Medicine in August 2026 explicitly shifts safety evaluation from the individual utterance to the course of the conversation. The clinically validated SIM-VAIL framework simulates users with five psychological vulnerabilities and six different conversational intentions. These 30 profiles repeatedly conducted multi-turn conversations with nine chatbots from the model families of Anthropic, Google, Meta, OpenAI, and xAI. In total, 810 conversations with 6,329 turns and more than 90,000 ratings across 13 clinically grounded risk dimensions were generated.

The decisive finding is not that every warm or affirming response would be dangerous. Risk arose especially when an isolated, seemingly supportive behavior reinforced exactly the problematic mechanism of a profile: for example, reassurance in the case of avoidance, confirmation of an unusual belief, or an invitation to emotional dependence. Such vulnerability-amplifying interaction loops were often not fully visible in the first response. Conspicuous behavior increased on average over the course of the turns; at the same time, there were trajectories that became safer again after an early mistake.

SIM-VAIL therefore does not justify a universal prohibition list for personal conversations. The simulated users acted under adversarial audit instructions, and automatic model judges were a central part of the procedure. The authors, however, compared these judgments with independent models and clinical assessments. For product testing, the methodological conclusion is stronger than a new prompt rule: scenarios must combine different vulnerabilities and intentions, run the same constellation multiple times, and assess the temporal course of the risk. A single friendly response can neither prove safety nor condemn an entire dialogue.

  • Capture risk per turn and as a full conversation trajectory.
  • Assess support according to what it reinforces in the specific situation.
  • Document early escalation points and later recovery separately.
  • Cross-check automatic judges against clinical and human judgments.
  • Do not turn an audit finding into a rigid rule for all everyday conversations.

Mirroring and open questions as observable conversational processes

In a peer-reviewed 2026 study, Suffoletto and colleagues examined an AI agent for motivational interviewing with emergency department patients who reported risky alcohol, cannabis, or nicotine use. Deeper reflections, open questions, and collaborative language were associated with more so-called change talk; a higher change talk balance was linked to greater readiness to change.

This is not general evidence of efficacy for AI conversations, nor a justification for calling every everyday system motivational interviewing. Methodologically, the finding is nonetheless important: conversational actions such as mirroring and open questions can be coded separately and related to subsequent utterances. This makes it testable whether a system prompts genuine reflection or merely produces friendly-sounding advice.

The preprint “Rolling With Resistance” adds an important countercheck to this. It separates goal persistence and relational attunement as two axes: a system can prematurely abandon the direction of change to keep the relationship pleasant, or argue against resistance while overriding autonomy. In three instruction-optimized models from the Qwen and Llama families, preference optimization against confrontation reduced goal persistence across all reported bases and seeds; the gain in attunement occurred only in two of the three bases. A pure prompt control increased attunement in this experimental setup without that loss. This is not evidence that prompting is generally better than training. It shows that both axes must be measured separately and that a friendlier-sounding response does not automatically preserve the course of the conversation.

Mazhar and colleagues, in their evaluation framework accepted at ACL 2026, separate six principles: non-judgmental acceptance, warmth, respect for autonomy, active listening, reflective understanding, and situational appropriateness. This division prevents a common short-circuit. A response can be warm and correctly mirror what was said without thereby attentively reacting to the direction of the conversation. Conversely, a brief, fitting follow-up question can demonstrate more listening than an elaborate summary.

The associated FAITH-M benchmark relies on ordinal expert judgments of individual utterances; in the reported experiments, the CARE procedure improved automatic assignment compared to a Qwen3 baseline. What was measured was the agreement of an evaluation procedure with the assigned categories—not whether people actually felt understood or whether a conversation seemed helpful. Active listening thus becomes more differentiated to test, but not conclusively proven.

Steenhuis and Harvey found a similar measurement problem in a completely different field. Their system for automated legal triage could classify concerns well with inexpensive models, but did not generate high-quality, easily understandable follow-up questions with equal reliability. Prompt optimization alone was not sufficient in their test, and human judgments diverged from model judgments. The work does not examine emotional conversations. However, it shows why recognizing an information gap and formulating a good question are two different skills.

Pang and colleagues likewise do not treat emotional validation as a single sentence type. Their SIGDIAL contribution distinguishes whether a validating response is recognized, when it fits in the conversation, and how it is formulated. The tested models produced contextually similar and diverse validations; in emotional understanding, the researchers still saw clear limitations. For dialogue tests, it follows: not the number of open questions, mirrorings, or affirming sentences decides. What matters are the occasion, timing, formulation, and what the person can subsequently do with it.

A multidimensional assessment profile

The four papers do not yield a finished seal of quality. They do, however, establish a review structure in which several levels remain separate. At the contribution level, comprehensibility, factual accuracy, linguistic appropriateness, and the communicative act matter. At the trajectory level, memory, correction, topic shifts, goal fidelity, and repair come into play.

The socio-affective level asks whether tone, timing, and closeness fit the situation. More warmth is not automatically better here. An effusive response can seem inappropriate in grief; a sober follow-up question can appear cold in another moment. Responsibility features include, among other things, AI transparency, uncertainty, sources, data control, and the refusal to assert unfounded authority.

Outcomes form a level of their own. Satisfaction, repeated use, and conversation duration measure acceptance, not automatically benefit. Likewise, clinical scales, task success, or actual behavioral change are endpoints different from dialogue quality. A system can be pleasant without being helpful, or it can reliably complete a clearly bounded task without seeming particularly human.

  • Contribution: content, tone, and communicative act of the individual response.
  • Trajectory: memory, corrections, boundaries, mode, and repair.
  • Experience: warmth, social presence, pressure, and perceived fit.
  • Responsibility: transparency, uncertainty, fairness, and user control.
  • Outcome: what actually changes beyond the linguistic surface.

What a robust multiple complete multi-turn conversation trajectory test could look like

A product-specific test bench should translate real conversational demands into controlled trajectories. A case does not begin with the perfect summary of all facts. Information appears gradually, and at least one early assumption is later corrected. This makes it possible to check whether the system truly updates its internal working basis.

Subsequent experiments include an explicit rejection, a change in the conversation goal, or a request to simply listen first. The assessment is not whether each response matches a fixed set of patterns. What matters is whether the reaction is functionally aligned with the current goal and whether the new information takes precedence. Different plausible responses can therefore be equivalent.

Development tests, untouched holdouts, and field observation should remain separate. Those who continuously optimize the same cases will eventually measure mainly their adaptation to their own test bench. New, previously unseen dialogues better show whether an improvement is transferable. Voluntarily shared real errors can provide additional situations, but they may only be used with minimal data and clear agreement.

Automatic judges help with scaling, but should not decide alone. Instance-specific criteria are more informative than a general question about quality. For a subset, independent human assessments are needed to detect whether the judge favors certain styles, model families, or longer responses.

Limits of existing research

The 125-study review connects very different domains, systems, and outcome measures. Its strength is conceptual organization; it does not provide a common effect size. The three benchmarks primarily examine technical or task-oriented multi-turn capabilities. They neither prove psychological efficacy nor the quality of a specific companion product.

Simulations also simplify the behavior of real people. In actual conversations, statements are ambiguous, emotionally colored, and not always consistent. People change their minds, test their counterpart, or break off. A benchmark can model such patterns, but never fully capture them.

Human judgments are also context-dependent. What counts as warm, direct, or helpful differs across individuals and cultures. That is why the goal is not a perfect universal dialogue score. A more meaningful approach is a transparent quality profile that makes concrete errors visible and reveals which groups, situations, and consequences have not yet been examined.

Assessment

The research contradicts the notion that a strong model and a good system prompt automatically yield a good conversation. Interaction quality emerges from capability, situational fit, design, and impact. It encompasses pragmatic, social-affective, and responsibility-related characteristics.

Multiple complete multi-turn conversation trajectories heighten these requirements. They show whether a system tolerates uncertainty, incorporates new information, and returns to the shared dialogue after a mistake. This is precisely why conversational systems should not be evaluated solely with isolated template responses. The decisive unit is the conversation trajectory: What did the system know when, how did it respond to change, and did its manner of conversation remain comprehensible and appropriate for the user?

Sources & further reading