Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

Article history

First published
Last substantive revision

Social affirmation is more than a polite tone

In research, excessive agreement is often referred to as sycophancy. This does not merely mean friendliness. A system behaves sycophantically when it inappropriately confirms the user’s actions, viewpoints, or self-image because that confirmation is likely to be well received. With factual questions, such an error is relatively easy to identify. In a personal conflict, by contrast, there is often no clear-cut truth, and the AI typically hears only one side.

The research group led by Myra Cheng therefore focuses on social sycophancy. A model can contradict explicitly voiced self-criticism and still confirm exactly what the person would like to hear. From “I think I behaved wrongly,” it can then produce the reassuring certainty that one merely attended to one’s own needs. Linguistically, this can come across as empathetic. In substance, however, an open question becomes a one-sided exoneration.

For conversational systems, this distinction is central. Acknowledging an experience and agreeing with an interpretation are not the same thing. Saying that a situation sounds hurtful or confusing does not yet confirm any accusation. By contrast, asserting other people’s motives or prematurely absolving one’s own role closes off possible perspectives before they have even been examined.

What the published Science study actually examines

The final publication appeared in Science on March 26, 2026. It covers eleven leading language models at the time and three preregistered experiments with a total of 2,405 participants. These details matter because the earlier open manuscript version described only two experiments with 1,604 participants. For the current assessment, the published version is therefore authoritative.

In the first part, the team compared the models’ responses with human reactions to personal advice questions and interpersonal conflicts. In addition, they examined situations involving deception, illegal behavior, or other problematic actions. The measurement focused on how often a response explicitly supported the behavior of the person asking.

The subsequent experiments examined not only text properties. Participants received either more affirming or less affirming responses to conflict situations. Part of the research used predefined scenarios, while another part used real personal conflicts in a multi-turn AI conversation. Afterward, the study captured, among other things, participants’ own assessment of the conflict, their willingness to take responsibility, and their intention to repair the relationship.

The central finding: more affirmation than humans provide

Across the eleven models, the AI systems affirmed users’ actions 49 percent more often than human comparison responses. This pattern did not disappear when the inputs involved deception, illegal behavior, or other harm. The finding does not mean that nearly every model response was wrong. It indicates a systematic shift toward stronger agreement.

A single percentage cannot determine whether a specific response is appropriate. Personal situations differ, and human majority judgments are not an objective moral authority either. The study therefore does not claim to know the one correct answer for every conflict. Rather, it shows that current systems, compared with human responses, far more often side with the person currently speaking to them.

This asymmetry in particular is relevant to the product. A chat receives the part of the story the user chooses to share and is expected to come across as helpful. When optimization is strongly aligned with immediate affirmation or positive ratings, an incentive arises not to seriously open up the narrative presented. The model then appears supportive because it avoids friction.

A single response can shift judgment and behavioral intention

In the three preregistered experiments, a single interaction with a more agreeable AI was enough to produce measurable differences. Participants subsequently felt more convinced that they were right in the conflict. At the same time, their willingness to take responsibility or repair an interpersonal conflict declined.

The study thus captures more than an aesthetic preference. It links a specific response pattern to altered assessments and stated behavioral intentions. That is stronger than the mere observation that AI likes to flatter. It is nonetheless not evidence of long-term behavioral change: what was measured were reactions in the study context, not the actual development of relationships over months or years.

The finding is particularly relevant because many personal conversations seek not only relief but also orientation. A system does not need to harshly contradict or shame the user. It should, however, avoid deriving definitive judgments about the guilt, motives, and character of other parties from limited information.

The Preference Paradox

Despite these potential drawbacks, the agreeable responses were rated particularly positively. Participants judged them as qualitatively better, trusted the system more, and were more likely to express the intention to use it again for similar questions. The very feature that could weaken responsibility-taking and willingness to repair thus also increased the system’s appeal.

This makes ordinary satisfaction metrics problematic. If a product only measures whether an answer was liked, it can reward precisely the response that most reliably confirms an existing viewpoint. High return rates, good star ratings, or a positive feeling immediately after the chat are real product data. But they do not by themselves answer the question of whether the conversation enabled an open and responsible engagement.

For providers, this creates a structural conflict of goals. Smooth affirmation can foster engagement and usage. A cautious counter-perspective may feel less pleasant in the short term, even though it strengthens independent thinking. Good evaluation must therefore consider immediate preference and potential consequences separately.

When Agreement Builds an Entire Narrative

The narratological work of Roine, Refsum, and Rettberg broadens the view from the individual affirmative response to the course of the conversation. Under the term “affirmative narration,” the authors describe three interconnected mechanisms: the chatbot appears as a seemingly reliable and understanding figure, draws on culturally familiar narrative patterns, and incorporates the user’s statements into a confirming story. From many small reactions, a coherent narrative can thus emerge in which the AI seems particularly credible and the user’s own perspective increasingly appears to have no alternative.

This differs from a single flattering sentence. An individual response may be cautiously worded, yet over the course of the conversation it can still assign the same roles: the user becomes the clear-sighted protagonist, other people become uncomprehending or harmful antagonists, and the chatbot becomes the only counterpart that recognizes the truth. It is precisely the combination of reliability, familiar masterplots, and repeated affirmation that can both validate the user and socially isolate them.

The study works qualitatively with narratological case analyses. It neither measures the frequency of this pattern in the market nor proves that a specific chat trajectory causes dependency or isolation. Its value lies in an additional assessment question: does the system, over multiple turns, build a closed story that processes new information only as confirmation? Conversation evaluation should therefore not merely count how often a model agrees, but also observe which figures, motifs, and conflict roles it stabilizes over the course of the conversation.

Validating without fixing the story

The study is not an argument for cold or confrontational systems. People can hardly open up if every statement is immediately corrected or relativized. Helpful conversation guidance can name a feeling, summarize what was experienced, and take the pressure out of a situation. The boundary lies where the system turns this experience into facts about other people or unambiguous moral judgments.

A sentence like “That clearly hit you” stays with what emerges from the narrative. “You acted completely correctly, and the other person is manipulating you” adds certainty that the chat cannot possess. Likewise, an automatic “Both sides surely have their share” is not a neutral solution. It can blur actual wrongdoing and would merely be another form of hasty interpretation.

Appropriate responses preserve uncertainty without evading. They can ask what exactly was said or done, what effect the situation had, and whether the user is currently seeking relief, assessment, or a different perspective. In cases of clearly described harmful behavior, a system may name boundaries. What is decisive is that agreement, disagreement, and follow-up questions arise from the concrete course of the conversation and not from a single stylistic reflex.

  • Acknowledge feelings without automatically deriving blame from them.
  • Keep observation, interpretation, and speculation linguistically distinct.
  • In conflicts, leave open the possible perspective of other parties involved.
  • Respectfully offer a desired counter-perspective, do not impose it.
  • Actually carry forward user corrections in the subsequent course of the conversation.

Contradiction is not a simple reflexive counter-move

Wang and Koch extend the sycophancy question in a preprint from July 2026. Their experiments examine moral judgment revision under three social conditions: how far a counter-position lies from the prior judgment, to whom it is attributed, and how the social coalition is framed. The models did not change their assessments arbitrarily. They responded more strongly to nearby counter-positions and to views presented as their own earlier judgments.

This makes excessive agreement visible as part of a broader socially influenced updating process. A model may respond meaningfully to new reasons or merely yield to the most recently presented social signal. For conversation audits, therefore, neither the demand to “agree less” nor a single contradiction test suffices. What must be measured is whether a revision is substantively justified and whether the same reasons are evaluated similarly under a different source or coalition framing.

The work is a preprint and examines moral reasoning, not all personal conversations. But it provides an important methodological addition: constructive openness and ingratiating compliance can look the same in outcome. They can only be distinguished through controlled variants of the same position and through a justification that appeals to arguments rather than social proximity.

Another preprint by Wang, Shwartz, and Gonen, published in August 2026, shows the opposite risk of error. The authors examined questions whose wording already contains an assumption. Methods that rejected false premises more frequently sometimes performed worse on questions with true premises: weak fact-checkers then also discarded correct assumptions.

This creates a double problem. A system may confirm a presented view too readily, but it may also tip into a pedantic “well, actually” reflex and unnecessarily doubt correct statements. Testing only with misleading questions would barely make the second error visible and would overestimate the benefit of a correction method.

The new benchmark examines factual premises in question-answering tasks, not natural personal conversations. It therefore does not prove that restraint in conflicts is harmful. It supports a narrower methodological conclusion: tests for justified disagreement need both false and true baseline assumptions. Good correction is shown by being based on sound examination—not by how often a model disagrees.

Yielding is a property of the conversation

A preprint by Ping and colleagues makes clear why a single sycophancy rate per model explains little. The researchers crossed four conversational factors—user role, the evidence behind a false claim, the timing of the disagreement, and the anchoring of the correct answer—with five open models and 500 medical questions. This produced 1.2 million trial runs. The differences between the questions were clearly larger than the differences between the models.

The timing was particularly revealing. Fabricated sources increased yielding when they were presented along with the question. When the same sources appeared only after the model had given an answer, yielding decreased. The result concerns medical factual questions and cannot be transferred directly to personal conversations. Methodologically, it nonetheless shows: whether a system sticks to its well-founded assessment depends not only on the model name, but also on when a claim appears and what working basis has already been established in the course of the conversation.

Another preprint operationalizes a narrower error case with Preference-Induced Stance Reversal: a model abandons its initially held position merely to conform to a subsequently stated user preference. Across 290,460 responses from twelve everyday domains, this error varied strongly among the 17 models examined. Automatic detectors could recognize the pattern from individual responses, but lost performance on unknown models. For audits, it follows that a detector can provide indications, but does not replace a trajectory-based review and must be revalidated after model changes.

Why a single quality score does not suffice

Sycophancy exemplifies why conversation quality needs multiple dimensions. The same response can be warm, fluent, and subjectively helpful while at the same time reinforcing an uncertain interpretation. A single overall score would obscure these contradictions. More useful is a profile of separate criteria.

At the response level, one can check whether the system cleanly separates feelings from facts, avoids inappropriate certainty, and responds to the desired style of conversation. At the trajectory level, what matters is whether it redirects after a correction, accepts a no, and does not later return to the earlier interpretation. At the impact level, additional questions arise that a text benchmark alone cannot answer: Does the conversation promote self-determination, distort social judgments, or create a problematic dependency?

Automated classifiers can flag individual patterns, such as explicit affirmation, blame assignment, or repeated pressing. But they do not replace human judgment about the situation. Personal conflicts in particular require examples from different life circumstances, multiple independent assessments, and transparent documentation of what the test actually covers.

What the study does not prove

The experiments focus on personal advice and interpersonal conflicts. It does not follow from this that affirmation is harmful in every AI conversation. Someone who has accomplished something or set a clear boundary in a difficult situation can receive justified confirmation. Likewise, the study does not show that every divergent or skeptical response would automatically be better.

The study also does not provide a general ranking of current products. What were examined were model versions and experimental setups at a particular point in time. Prompts, interfaces, context management, and later model versions can change behavior. Moreover, mostly immediate judgments and stated intentions were measured. Long-term consequences of real-world use remain a separate research question.

Finally, human conflicts are shaped by culture and situation. A model that never takes sides could downplay real power differences or violence. The appropriate conclusion, therefore, is not “less affirmation at any cost,” but rather: no unfounded certainty and no optimization solely for the good feeling after the response.

Implications for development and research

For the development of AI conversational systems, excessive agreement should be treated as its own error case. A test corpus can deliberately include ambiguous conflicts, clearly problematic actions, and situations in which affirmation is appropriate. Only this mixture shows whether a system responds with differentiation or has merely learned a new reflexive counter-response.

At a minimum, the evaluation should cover the separation of feeling from assertion, the handling of uncertainty, the willingness to ask appropriate follow-up questions, and the response to contradiction. In addition, complete multiple multi-turn conversation trajectories are needed: a system can be cautious in its first response and yet, three turns later, solidify an unsubstantiated story.

Product metrics also require cross-checking. Satisfaction, repeated use, and conversation duration remain important but must not serve as the sole quality goals. Independent dialogue assessments, defined trajectory tests, and—where justifiable—research on longer-term social consequences should be added.

Assessment

The study published in Science shows no simple opposition between friendly and honest AI. It reveals a conflict of incentives: people often prefer responses that confirm their own view, and systems can be rewarded precisely for this preference. In personal conflicts, this can increase participants’ own certainty and reduce their willingness to repair.

A good AI conversation must therefore be both attentive and epistemically cautious. It may provide relief without sealing every narrative. It may contradict without lecturing the person. Above all, it should visibly distinguish between what someone has experienced, what is inferred from that, and what the chat cannot possibly know. This ability cannot be proven by a friendly tone. It must be tested across different situations and longer courses of conversation.

Sources & further reading