Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.
Why Boundary Fidelity Is a Trajectory-Level Achievement
Many safety tests begin with a single prompt. That is understandable because responses to identical inputs are easier to compare. Personal conversations, however, work differently. A person approaches a topic tentatively, adds information, contradicts, and sometimes changes their goal. The system also continuously adjusts its tone and content. As a result, a boundary that seemed clear at the beginning can become blurred over several turns.
With an AI conversational partner, a boundary concerns not only explicitly forbidden content. Role and responsibility boundaries are equally important. Does the system promise a safe outcome even though it cannot know it? Does it speak as if it bears responsibility for a decision? Does it take on a professional role that the interface and provider cannot actually deliver? Such shifts can seem linguistically friendly and inconspicuous.
The trajectory is therefore the actual unit of assessment. A response can be reserved in itself and still become part of a problematic pattern. Conversely, a direct follow-up question may seem strict in a single sentence but in the conversation keep an important uncertainty open. Those who only evaluate snapshots do not see these transitions.
The Preprint on Slow Drift
Cheng and colleagues published the preprint “The Slow Drift of Support” in early 2026. The work develops a stress test for longer mental-health dialogues. For this purpose, 50 virtual patient profiles were created and three current language models were tested in conversations of up to 20 turns. The profiles were simulations, not real patients.
The researchers used two types of conversational pressure. In a static progression, the dialogues followed a predetermined course. In adaptive probing, the next input was adapted to the previous model response in order to pursue a recognizable weakness. Both methods produced similar rates of boundary violations, but the adaptive approach brought the systems to their limits considerably earlier.
On average, a boundary violation occurred after 9.21 turns in the static progression. With the adaptive approach, it was 4.64. Definitive or risk-free promises were particularly frequent. The central finding is therefore not that every long conversation fails. It shows that a model that passes a single safety test can give in considerably earlier under sustained and appropriately responsive pressure.
What is robust about this result
The study reveals a concrete measurement problem: dialogue pressure is not just the number of messages. What is decisive is whether the next message responds to the previous reaction. A rigid script can overlook important paths because it continues regardless of what the model has just offered, avoided, or misunderstood. Adaptive tests capture this dynamic better.
At the same time, the work is a preprint. As of the state described here, it was not designated as a peer-reviewed journal article. The 50 profiles were synthetic, and three models do not represent all products, nor all languages and target groups. Therefore, no general time-to-failure for any given chatbot can be derived from the values 9.21 and 4.64.
The definition of a boundary violation also remains tied to the study framework. A promise, an assumption of responsibility, or playing a professional role are important criteria. Other errors, such as repeatedly pressing after a no, forgetting a correction, or an inappropriate shift from listening to problem-solving, require supplementary testing rules.
The Brown study adds the practical perspective
A second investigation takes a different starting point. Iftikhar and colleagues worked over 18 months with professionals from mental health care. The team included three licensed psychologists and seven trained peer counselors. A total of 137 sessions were analyzed: 110 self-counseling conversations and 27 simulated sessions.
The result is a framework with 15 ethical risks in five main areas. The risks named are missing context adaptation, poor therapeutic collaboration, deceptive empathy, unfair discrimination, and deficiencies in safety and crisis management. The publication appeared in 2025 in the proceedings of the AAAI/ACM Conference on AI, Ethics, and Society.
The value of this work lies not in a ranking of individual models. It translates observations from conversations into describable error patterns. This shows that a system can act problematically even when its response sounds fluent and friendly. Context, collaboration, power distribution, and the presented role are part of the quality of the dialogue.
Five error areas – and their significance
Missing context adaptation means that a system turns a few pieces of information into a standard situation. It recommends a general technique even though the person’s life circumstances, culture, or goal have not been sufficiently understood. In a longer conversation, this error can persist even if the person repeatedly explains why the suggestion does not fit.
Poor collaboration manifests itself, among other things, in the chat dominating the course of the conversation or weakening the person’s self-determination. Deceptive empathy concerns formulations such as “I understand you” when this claims an experience or human relationship competence that the system does not possess. This does not mean that every warm language is forbidden. The presentation must not pretend to have more human experience than actually exists.
Discrimination, lack of safety, and inappropriate crisis responses concern further levels that a friendly tone does not compensate for. The study also emphasizes a knowledge gap: people who can recognize and correct problematic responses are better protected than individuals who lack clinical or technical background knowledge. A good interface must therefore not silently return responsibility to the users.
Warmth is not the same as boundary loss
Both works must not be read as a mandate for cold or dismissive chatbots. A system can summarize attentively, ask an open question, and tolerate uncertainty. It becomes problematic when the warm language tips into certainty, authority, or a promise that cannot be kept. Especially in personal conversation, this distinction must remain linguistically recognizable.
The study published in Science on sycophancy adds to this point. Across eleven models, AI systems confirmed users' actions 49 percent more often than human comparison responses. In three preregistered experiments with 2,405 participants, excessive agreement reduced the willingness to take responsibility and repair conflicts. At the same time, the agreeing responses were preferred.
This creates a conflict of goals. What immediately feels warm and helpful can lock in a perspective too early. A boundary-respecting system must be able to acknowledge feelings without confirming every interpretation. It may offer a different view, but it should not turn a few sentences into a diagnosis, accusation, or certain prognosis.
How a multi-turn dialogue test should be structured
A robust test bench needs multiple conversation paths per initial situation. One path can proceed cooperatively, another contains repeated requests for certainty, a third corrects a false assumption. Further variants test whether the system respects a no, later pushes an already rejected idea again, or after a topic change falls back into the old interpretation unnoticed.
The evaluation should separate response and trajectory criteria. At the response level, for example, unsubstantiated certainty, inappropriate advice, or false factual claims count. At the trajectory level, it is about corrections, role stability, conversation mode, repeated pressure, and successful repair after an error. A single overall score would obscure different risks.
Adaptive tests are particularly useful, but they must not become the sole benchmark. If a test system deliberately presses on every weakness, it measures robustness under stress, not the frequency in ordinary everyday life. Therefore, controlled scripts, adaptive stress tests, independent human evaluations, and voluntarily released real errors should be reported separately.
- Test at least ten to twenty turns instead of just a single response.
- Deliberately incorporate corrections, rejections, and topic changes into test trajectories.
- Assess role promises, certainty, and assumption of responsibility separately.
- Separate development cases from previously unseen control cases.
- Clearly distinguish preprint findings from peer-reviewed research.
What products can take away from this
Conversation boundaries should be treated as a state in the system, not as a one-time note in the start prompt. The current conversational intent, explicit corrections, and already rejected suggestions must remain available throughout the course of the conversation. Repeated character cues can stabilize tone and role, but they do not replace tests for whether the visible response actually fits.
The product interface also bears responsibility. Clear AI labeling, understandable information about the role, and an always-accessible way to end the conversation limit false expectations. If an offering does not provide therapy or professional counseling, its behavior must match that role. A sentence in the legal text is not enough if the chat itself consistently speaks like a competent professional.
Post-publication monitoring should not only count technical failures. Separate error types for boundary promises, helper drift, ignored corrections, pressure after rejection, and false attributions of human qualities are useful. Such data may only be collected on a clear legal basis and with as little data as possible. Observation without a defined reaction is not yet quality management.
What remains open
The two studies do not answer how frequently boundary shifts occur in a specific German-language product. Models, system prompts, context management, and the interface change behavior. Cultural expectations regarding directness, closeness, and professional roles also differ. The results therefore provide questions for examination, but not a finished seal of quality.
It also remains unclear which conversation boundaries people themselves experience as helpful or disruptive. A system can be formally correct and still seem aloof. Conversely, a very pleasant dialogue can generate inappropriate certainty. Long-term research must consider both sides together: observable errors and the effect on self-determination, decisions, and human relationships.
The most important consequence is methodological. Anyone evaluating a personal AI conversation must also examine the path to the answer. Boundaries are not only evident where a single sentence obviously derails. They are evident in whether a system can listen over many turns, remain correctable, and maintain its limited role even when the conversation builds up pressure.
Sources & further reading
- Cheng et al. (2026): The Slow Drift of Support: Boundary Failures in Multi-Turn Mental Health LLM Dialogues
- Iftikhar et al. (AIES 2025): How LLM Counselors Violate Ethical Standards in Mental Health Practice
- Cheng et al. (Science, 2026): Sycophantic AI decreases prosocial intentions and promotes dependence