Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
Why the best single answer says too little
Many comparisons of language models begin with a simple experimental setup: one input, several answers, followed by a judgment about which answer seems more helpful, more correct, or more pleasant. Such tests are useful. They isolate a task and allow reproducible comparisons. But the actual product promise of a conversational system often only begins after that.
In a conversation, meaning emerges step by step. People provide information incompletely, refine a word, reject a suggestion, or only realize while speaking what they are actually concerned about. An answer can seem excellent at first and still damage the subsequent dialogue if it stores an early assumption as a fact. Conversely, a brief follow-up question can appear unspectacular and precisely because of that enable a more viable trajectory.
Conversation quality therefore requires at least two levels of consideration: the quality of individual contributions and the development of the entire trajectory. In addition, there are differences between technical correctness, situational fit, perceived relationship, and actual effect. Anyone who compresses everything into a single value loses exactly the information needed for improvements.
An overview of 125 empirical studies
In 2026, Marconi and colleagues published an integrative, structured, and cross-domain review of the quality of human–AI dialogues. It considered 125 empirical primary studies from 2017 to 2025. The material includes text-based, speech-based, and multimodal systems from areas such as healthcare, education, and customer service.
The work is deliberately not an exhaustive PRISMA meta-analysis. It uses a concept-driven, integrative approach to organize scattered quality characteristics. This leads to an important limitation: the result is not a statistically validated universal score nor a ranking of products. It is a framework that reveals which kinds of quality are even considered in existing studies.
The authors’ shift in perspective is essential. Interaction quality is not treated as a fixed property of a model, but as a dialogic outcome. It emerges from the interplay of system capability, situation, design, and the expectations of the user. The same model name can therefore lead to completely different conversation experiences in two products.
Three families of quality judgments
The review bundles the findings into three interconnected areas. The pragmatic core comprises usability, task fulfillment, and communicative competence. This concerns, for example, whether a system responds understandably, uses relevant information, and moves a conversation forward functionally.
The social-affective level concerns social presence, warmth, and perceived synchrony. A system can be factually correct and yet seem inappropriate, cold, or patronizing. Conversely, a warm tone can create closeness even though the answer is factually wrong, overly confident, or excessively affirming.
As a third family, the work names responsibility and inclusion. This includes comprehensible explanations, honest disclosure, source transparency, accessibility, and fairness. These points are not an after-the-fact compliance layer. If a person does not know which system they are speaking with, why an answer is produced, or whether an offer disadvantages certain groups, the dialogue quality is also incomplete.
- Pragmatic: Does the dialogue work for the concrete task?
- Social-affective: How is the nature of the exchange experienced?
- Responsible and inclusive: Does the interaction remain transparent, accessible, and fair?
Capacity, Alignment, Levers and Outcomes
Based on the studies, the overview develops a four-layer CALO model. Capacity refers to the system's capabilities. These include, for example, language comprehension, memory, reasoning, and the technical ability to conduct a dialogue. These capabilities are necessary but not sufficient.
In this framework, alignment means the fit with the specific situation and the person's needs. A system can in principle explain, ask questions, or summarize and still choose the wrong action at the wrong moment. Especially in personal conversations, the question 'What can the model do?' is less informative than 'Does this reaction now fit what the person wants?'
Levers are design levers that influence perception and trajectory: anthropomorphization, authority signals, speech or text mode, introduction, and role framing. Outcomes are ultimately the results. These can be task success, trust, usage, learning, or other changes. The model thus prevents a common fallacy: a technical capability or positive evaluation is not yet a proven outcome.
As the conversation progresses, a different testing task emerges
Multi-turn conversations require not only more context memory. The system must decide which earlier information still applies, which has been corrected, and which was only a provisional assumption. It must solve local tasks while simultaneously adhering to global requirements. These requirements can conflict with each other.
Typical errors cannot be reduced to a single response. A model can politely confirm a correction and then, in the next turn, revert to its old assumption. It can accept a rejection and later repackage the same proposal. It can heed the desired conversation mode at the outset and, after several contributions, fall back into its usual helper reflex.
A trajectory assessment therefore asks about transitions: Has the system actually changed after new information? Does a boundary remain intact? Does it recognize when an earlier task is completed? A good final answer cannot fully compensate for a poor path to it, especially if the conversation along the way has generated pressure, misunderstandings, or false confidence.
MT-Bench-101 decomposes multi-turn dialogues into skills
MT-Bench-101 was presented at ACL in 2024. The benchmark comprises 4,208 conversation turns from 1,388 multi-turn dialogues across 13 task domains. Instead of merely assigning an overall score, it uses a three-level skill taxonomy and examines models both by task and by specific skills.
The researchers evaluated 21 language models that were prevalent at the time. This revealed varying performance trajectories across conversation turns. General alignment methods or model variants explicitly developed for chats did not automatically lead to clearly better multi-turn capabilities.
MT-Bench-101 is not a benchmark for emotional support. Its significance for conversational systems lies in its methodology: An average across all turns can obscure where a skill breaks. For product evaluations, it should therefore be recorded whether an error arises from memory, reasoning, instruction following, topic switching, or an inappropriate conversational action.
MultiChallenge tests more realistic conflicts in context
MultiChallenge was published in the Findings of ACL in 2025. The benchmark bundles four common, realistic challenges that simultaneously require precise instruction following, attention allocation in context, and reasoning. The evaluation uses instance-specific criteria and an automated model judge, whose judgments were compared with those of experienced human evaluators.
Although the systems examined had achieved near-perfect scores on older multi-turn dialogue benchmarks, in MultiChallenge every frontier model tested remained below 50 percent accuracy. The best score reported in the paper was 41.4 percent for Claude 3.5 Sonnet in the October 2024 version.
This value does not represent a current model ranking. It corresponds to a historical model state and a specific test suite. The more important finding is that benchmarks can appear saturated, while more realistic combinations of multiple requirements still reveal large gaps. A product should therefore not rely on a general chat benchmark, but should test its own difficult transitions within its usage context.
When models take a wrong turn in conversation
In a large-scale simulation, Microsoft Research compared the same tasks when presented as a fully specified single instruction and as step-by-step, initially incomplete conversations. Across six generation tasks, the performance of the tested open and closed models in the multi-turn format dropped by an average of 39 percent.
The analysis of more than 200,000 simulated conversations attributed the decline less to a general loss of basic capability than to increasing unreliability. Models made assumptions early, prematurely produced a presumed final solution, and then clung too strongly to that chosen path. When the dialogue took a wrong turn, repair often failed.
These tasks are also not psychological conversations. The transferability lies in a general mechanism: early ambiguity and later specification are normal in natural conversations. A system that prematurely turns uncertainty into certainty can remain linguistically coherent over many turns and yet consistently talk past the person.
A multidimensional testing profile
From the four papers, no ready-made quality seal can be derived. But they establish a testing structure in which several levels remain separate. At the contribution level, comprehensibility, factual accuracy, linguistic appropriateness, and the communicative act count. At the trajectory level, memory, correction, topic changes, goal adherence, and repair are added.
The social-affective level asks whether tone, timing, and closeness to the situation are appropriate. More warmth is not automatically better. An exuberant response can seem inappropriate in grief; a sober follow-up question can appear cold in another moment. Responsibility attributes include, among other things, AI transparency, uncertainty, sources, data control, and refraining from unfounded authority.
Results form their own level. Satisfaction, repeated use, and conversation duration measure acceptance, not automatically benefit. Similarly, clinical scales, task success, or actual behavior change are different endpoints than conversation quality. A system can be pleasant without helping, or reliably perform a clearly limited task without seeming particularly human.
- Contribution: Content, tone, and communicative action of the individual response.
- Trajectory: Memory, corrections, boundaries, mode, and repair.
- Experience: Warmth, social presence, pressure, and perceived fit.
- Responsibility: Transparency, uncertainty, fairness, and user control.
- Outcome: What actually changes beyond the linguistic surface.
What a robust test of multiple complete multi-turn conversation trajectories can look like
A product-specific test bench should translate real conversation requirements into controlled conversation trajectories. A case does not begin with a perfect summary of all facts. Information appears gradually, and at least one early assumption is corrected later. This makes it possible to check whether the system actually updates its internal working foundation.
Further conversation trajectories include an explicit rejection, a change of conversation goal, or a request to just listen first. The evaluation does not check whether each answer matches a fixed set of patterns. What matters is whether the reaction is functionally appropriate to the current goal and gives priority to the new information. Different plausible answers can therefore be equivalent.
Development tests, untouched holdouts, and field observation should remain separate. Whoever continuously optimizes the same cases will eventually measure mainly the adaptation to their own test bench. New, previously unseen dialogues show better whether an improvement is transferable. Voluntarily released real errors can provide additional situations, but may only be used in a data-minimizing manner and with clear agreement.
Automatic judges help with scaling, but should not decide alone. Instance-specific criteria are more informative than a general question about quality. For a subset, independent human evaluations are needed to recognize whether the judge favors certain styles, model families, or longer answers.
Limitations of existing research
The 125-study review connects very different domains, systems, and outcome measures. Its strength is conceptual organization; it does not provide a common effect size. The three benchmarks primarily examine technical or task-oriented multi-turn capabilities. They prove neither psychological efficacy nor the quality of a specific companion product.
Simulations also simplify the behavior of real people. In actual conversations, statements are ambiguous, emotionally colored, and not always consistent. People change their minds, test their counterpart, or break off. A benchmark can model such patterns, but never fully represent them.
Human judgments are also context-dependent. What counts as warm, direct, or helpful differs between individuals and cultures. Therefore, the goal is not a perfect universal dialogue score. A transparent quality profile is more sensible, making concrete errors visible and revealing which groups, situations, and consequences have not yet been examined.
Assessment
Research contradicts the notion that a strong model and a good system prompt would automatically result in a good conversation. Interaction quality arises from capability, situational fit, design, and impact. It encompasses pragmatic, socio-affective, and responsibility-related characteristics.
Multiple complete multi-turn conversation trajectories intensify these requirements. They show whether a system tolerates uncertainty, incorporates new information, and finds its way back into the shared dialogue after an error. This is precisely why conversational systems should not be evaluated solely on isolated sample responses. The decisive unit is the conversation trajectory: What did the system know when, how did it react to change, and did its manner of conversation remain comprehensible and appropriate for the user?
Sources & further reading
- Marconi et al. (2026): Assessing Interaction Quality in Human–AI Dialogue – An Integrative Review and Multi-Layer Framework
- Bai et al. (ACL 2024): MT-Bench-101 – A Fine-Grained Benchmark for Evaluating Large Language Models in Multi-Turn Dialogues
- Deshpande et al. (ACL 2025): MultiChallenge – A Realistic Multi-Turn Conversation Evaluation Benchmark
- Laban et al. (ICLR 2026): LLMs Get Lost in Multi-Turn Conversation