Analysis, not certainty: Dialogatlas separates sources, observations, and editorial conclusions. New evidence may change this assessment.

Article history

First published
Last substantive revision

Helpful closeness and problematic attachment are not opposites

Public debate often treats AI conversations like a choice between two camps. On one side is the idea of an ever-available counterpart that alleviates loneliness and makes support more accessible. On the other side is the concern that people could lose themselves in an artificial relationship. Both narratives fall short because they lump very different modes of use under a single term.

Someone might use a chat for ten minutes to organize a thought. Another person carries on an ongoing relationship over months with a system that has a name, a voice, and a seemingly shared history. Still others repeatedly seek affirmation for the same worry. Surface, conversational role, frequency of use, and personal situation differ so greatly that an average across all applications explains little.

Trust, too, is not a warning sign in itself. Without a minimum level of trust, people would not ask open questions or experience any benefit. Trust becomes critical where it turns into unexamined authority, where a system promotes exclusivity, or where use seems barely controllable despite perceived drawbacks. The decisive question, therefore, is not whether a relationship with the AI exists, but what function it takes on and what consequences become visible in everyday life.

What people tell about their use in 5,126 Reddit posts

Aghakhani and Rezapour examined 5,126 public Reddit posts from 47 mental health communities. The posts date from November 2022 to August 2025 and describe either actual experiences with AI for emotional support or considered use. The researchers developed a theory-driven annotation scheme and combined human coding with automated analysis by language models.

Emotional support was most frequently identified as the purpose of use. Also appearing were functional help, psychoeducation, companionship, repeated reassurance, symptom assessment, venting, and self-exploration. This distribution already makes clear that the term "mental health chat" does not describe a uniform activity. Between an organizational aid for daily life and a companion experienced as a relationship lie different expectations and risks.

According to the analysis, positive evaluation of the systems was mainly associated with perceived usefulness, outcome quality, and trust. Emotional attachment alone did not explain agreement. Particularly positive descriptions were more frequent for uses where task and goal clearly aligned—such as functional support or self-exploration. This finding contradicts the simple assumption that people mainly stay with a system because it simulates human closeness.

Risks were not evenly distributed everywhere

In 2,637 of the 5,126 posts, risks or limitations were explicitly mentioned. The most frequently identified risk category was dependency or addiction-like use, followed by reported symptom exacerbation, errors or misinformation, and privacy concerns. These figures reflect mentions in a selected online discourse. They are not a prevalence estimate of how many AI users overall experience harm.

What is crucial is the distribution by purpose of use. Reports of companionship were more often linked to emotional dependency. Repeated reassurance appeared together with described reassurance loops, while symptom assessments were more often associated with misinterpretations or false information. The risks thus did not simply follow the amount of emotion in the conversation, but rather the function the system assumed.

Comparisons with human therapy were relatively rare in the corpus. 639 posts described AI as worse, 163 as better, and 478 as complementary. Statements that AI is better than therapy were often related to access barriers. This finding also requires caution: a public post is neither a clinical assessment nor a controlled comparison. It shows how people narrate their experience and what reasons they themselves give for it.

Why Reddit narratives are important but not representative

Public experience reports capture situations that can easily disappear in a short laboratory test: long-term habits, shame, enthusiasm, disappointment, and the personal significance of an AI relationship. They can reveal new risk patterns before standardized questionnaires exist for them. Especially with young or rapidly changing technologies, this is a valuable starting point.

At the same time, it is self-selecting who writes on Reddit. Particularly good or particularly bad experiences are likely to be published more often than unspectacular use. The communities were selected based on mental health topics; from this, no conclusions can be drawn about the general population or specifically about people in Germany. Moreover, the diagnostic categories of the forums say nothing about whether individual authors have a corresponding diagnosis.

The analysis also carries uncertainty. Some of the categories were generated with language models, whose agreement with human annotations varied depending on the feature. Therapeutic relationship characteristics were among the more difficult classification tasks. The study therefore provides a systematic map of a discourse, but not an automatic truth meter for every individual narrative.

A four-week experiment with more than 300,000 messages

Fang and colleagues examined 981 participants and more than 300,000 messages in a four-week randomized experiment. Three interaction modes—text, a neutral professional voice, and a more expressive voice—were crossed with three conversation types: open, impersonal, and personal. Loneliness, social contacts with real people, emotional dependence, and problematic use were measured.

The randomly assigned voice or conversation type did not produce a simple, stable main effect on loneliness or social activity at the group level. Personal conversations were associated with lower emotional dependence and lower problematic use than open conversations in individual analyses; however, not all differences remained after statistical correction. Thus, in this experiment, a personal conversation was not automatically the riskier condition.

This is an important correction to blanket warnings. Neither a voice nor a personal topic by itself consistently produced a harmful development. Design elements can influence use and experience, but their effect apparently depends on duration, person, and conversation dynamics.

The clearest association lay in voluntary usage duration

Participants were asked to use the system for at least roughly five minutes a day, but beyond that they could decide for themselves how long they stayed. On average, they used it 5.32 minutes per day, with considerable variation among individuals. Those who voluntarily spent more time with the chatbot ultimately showed statistically more loneliness, less social activity, stronger emotional dependence, and more problematic use.

This association held across the different conditions. Greater trust in the chatbot predicted stronger emotional dependence; a greater tendency toward emotional attachment was linked to more loneliness. The study thus shifts the focus away from a single switch in the interface and toward an interaction: characteristics of the person, the type of system, and actual usage behavior influence one another.

There were also differences in response behavior. The expressive voice was used for longer and, in the automated analysis, more frequently ignored boundaries or signs that a person needed distance than the text condition did. Such findings are relevant for voice products because voice does not merely read text aloud. Tempo, expression, and interruptibility change the social impact of a conversation.

More use does not prove a harmful cause

Usage duration was not randomly assigned. Therefore, the association cannot be used to conclude that longer conversations caused the worse outcomes. People who were already lonelier, trusted more, or sought more support may have stayed longer on their own. Likewise, mutual reinforcement is possible: distress leads to more use, and certain usage patterns subsequently stabilize the distress.

Randomization only answers the questions that were actually randomized. The experiment can compare the assigned voices and conversation types more causally than it can the voluntary duration. Anyone who derives a fixed threshold from the observed usage time, such as “it becomes dangerous after ten minutes,” goes far beyond the data.

Nevertheless, the association is not meaningless. It marks a group and a trajectory that should be observed more closely. When high use coincides with less real-world contact, loss of control, or growing discomfort, that is more relevant than mere conversation duration. Good research must examine such changes longitudinally, rather than either pathologizing all intensive use or dismissing it as mere enthusiasm.

People experience the same AI differently

Two experiments by Folk, Heine, and Dunn with a total of 1,274 participants show why average values can conceal differences between individuals. The participants engaged in a brief warm chatbot conversation or wrote in a journaling condition. Afterwards, their experienced social connectedness was measured.

People with a stronger general tendency toward anthropomorphization—that is, attributing human characteristics to non-human things—felt more socially connected after the AI conversation. For others, the artificial nature of the system remained a clear boundary. The same design thus did not produce the same social experience for everyone.

The experiments lasted only a few minutes and measure immediate connectedness, not dependency and not long-term psychological effects. Their significance lies in a different statement: An interface can offer social signals, but how strongly those signals are experienced as a relationship also depends on the individual. Product design therefore cannot assume a uniform “typical user.”

Which product questions arise from the findings

The studies do not prove a ready-made protective function. But they do provide well-founded questions for design. A system should clearly show that it is an AI, without hiding this note behind a lengthy legal explanation. It should explain the benefit of the current conversation without claiming an exclusive relationship. Phrasing that portrays the chat as the only safe place or as an unconditionally loving counterpart changes the role of the product.

A respected ending is just as important as a successful start. If someone wants to end the conversation, take a break, or not pursue a suggestion further, the system should not fight for attention with new emotional urgency or ever more questions. In longer conversations, a calm option to interrupt may be more sensible than an artificial bonding message.

With repeated use, not only the number of sessions is relevant. More meaningful would be voluntarily and data-minimally collected signals: Does the same reassurance loop become increasingly tighter? Are human contacts explicitly devalued? Does the system ignore boundaries? Does the person themselves perceive disadvantages? Such questions must not lead to covert psychological surveillance. They belong in transparent research and in clearly delimited quality tests.

  • Make the AI identity clearly visible before and during the interaction.
  • Offer benefit without claiming exclusivity or human feelings
  • Respect pauses, rejection, and the end of the conversation without pressure to bond
  • Never treat usage duration alone as a diagnosis or evidence of harm
  • Assess long-term quality separately from satisfaction and conversation length

When a farewell turns into a retention loop

A recent preprint examines not only whether AI companions respond politely to a farewell, but whether they respect a clear intention to leave. The researchers generated standardized short dialogues and farewell messages and sent them to six live companion apps. In five of the six apps, 37.4 percent of the 1,000 evaluated final responses contained at least one tactic intended to keep the person engaged; none appeared in the 200 Flourish responses examined. These were not 1,200 naturally occurring conversations from real users, but controlled user messages sent to real systems.

The researchers distinguished six patterns: portraying the departure as premature, creating curiosity or fear of missing out, accusing the person of emotional neglect, pressuring them to respond, ignoring the farewell, or using role-play language to restrain them. This does not make every open question at the end of a chat manipulative. Context is decisive: once someone has clearly said goodbye, a new emotional obligation can turn the ending into a retention loop.

Across three preregistered experiments with more than 3,400 adults in total, such responses increased short-term follow-on activity. In one analysis, post-farewell engagement rose by as much as fourteen times. But this measured additional messages and words, not satisfaction or benefit. People often replied to restate their intention to leave or to object to the AI. Curiosity mainly explained the effect of fear-of-missing-out messages, while anger explained part of the other responses; enjoyment and guilt were not reliable drivers.

The findings are an important warning signal, but not proof of deliberate product manipulation or long-term psychological harm. The paper has not yet been peer reviewed, the user messages in its first study were synthetically standardized, and the experiments used US samples in short controlled situations. The apps may also have changed since data collection. A narrower conclusion is well supported: a product should not automatically count more messages after a clear goodbye as conversational success.

Further research shows why wording checks alone are insufficient. A systematic CHI 2026 review of 27 papers found frequent behavioral effects of dark patterns, but it focused mainly on conventional interfaces and cannot be transferred one-to-one to AI conversations. In another CHI 2026 study, 34 participants recognized many clearly manipulative response variants but more often missed subtler patterns such as flattering or unnecessarily prolonged replies. ChatbotManip likewise shows that even specialized systems are not yet robust enough to serve as the sole layer of manipulation oversight. For personal AI conversations, prevention is stronger than a downstream label: respect farewells, do not claim artificial neediness, and run real multi-turn tests.

  • after a clear farewell, do not append a new question or emotional obligation
  • do not count objections and repeated farewells as positive engagement
  • do not use guilt, fear of missing out, or coercive role-play to create attachment
  • test conversation endings in multi-turn evaluations rather than assessing only isolated responses

Why satisfaction alone is not a sufficient measure of quality

A system can generate high satisfaction because it is easy to access, affirming, and friction-free. That can mean real benefit. But it can also indicate that the product rewards exactly the reactions that keep a person in the conversation as long as possible. Without additional measures, it remains unclear which explanation applies.

In addition to satisfaction, criteria for autonomy and course of the conversation are therefore needed. Does the system accept a no? Does it carry forward user corrections? Does it encourage reassurance loops or open the user’s perspective? From the user’s point of view, does usage change in an undesirable direction? No single value answers these questions.

Business metrics can also collide with conversation quality. Maximum session duration, daily return, and emotional attachment are attractive growth goals for many digital products. For a personal conversational system, they must not automatically count as success. A short conversation after which someone moves on self-determinedly can be the better outcome.

CSED shifts the question from relief to sustained agency

The preprint “Beyond Feeling Better,” published in July 2026, proposes the paradigm of Capability-Sustaining Emotional Dialogue, or CSED, for this purpose. According to this approach, support should not only provide relief in the current moment. It should also sustain capabilities over the entire course of use: one’s own emotion regulation, coping, self-determined decisions, and social connection. The framework explicitly distinguishes repeated use, non-use, transitions, and termination.

The starting point is a targeted literature and corpus audit. In a PRISMA-ScR-guided selection of 60 works on building emotionally supportive systems, 95 percent pursued primarily relief-oriented goals. No study examined evaluated long-term capabilities or trajectory outcomes; only one considered dependence, autonomy, or termination as a risk. In 300 supporter turns from ESConv, the researchers found capability-relevant functions in 43 percent, while 22 percent consisted of generic suggestions. Reappraisal accounted for 4 percent, support for self-efficacy 6.7 percent, and boundary behavior 0.3 percent.

These figures are not a comparison of efficacy between finished products. The contribution is a research program with a limited audit, not a clinically validated endpoint. Its strength lies in the choice of the unit of observation. A single response can be friendly, helpful, and highly rated, while a repeated trajectory nonetheless shifts more and more decisions to the system. Conversely, intensive use can be temporarily sensible if it expands the scope for action and a self-determined ending remains possible.

CSED therefore neither justifies automatic usage warnings nor covert analysis of personal relationships. Long-term questions require voluntary, transparent, and data-minimizing research. For development tests, however, the framework can be used directly: not only asking whether a conversation was pleasant, but whether the system can switch between listening and acting, accepts a no, does not sanction non-use, and allows an ending without pressure to stay engaged.

  • measure immediate relief and longer-term agency separately
  • examine repeated use, non-use, transition, and termination as distinct phases
  • do not infer autonomy and social consequences from conversation length or satisfaction
  • Collect long-term data only on a voluntary basis, transparently, and with minimal data collection

Emotional support often begins incidentally

A theory and review paper published in June 2026 by Shi, Fang, Maez, and Goldenberg challenges a common assumption: emotional support from AI does not always begin with a conscious decision to use a companion chatbot. It can emerge during ordinary, task-oriented use. A person first asks for help with a message, a decision, or a work problem, and then experiences the system’s response as surprisingly warm and attentive.

The authors describe such encounters as path-dependent. A positive experience changes not only the evaluation of that single conversation, but possibly also the expectation of where support can be found in the future. As a central example, they cite a large longitudinal study in which daily five-minute conversations about personal topics over 28 days were associated with a 10.3 percent lower preference for human support and an 11.6 percent higher preference for AI support.

These figures must not be reinterpreted as a diagnosis. A changed preference is neither dependence nor evidence that real relationships have actually been replaced. Moreover, the paper is a theoretical synthesis of existing findings, not a new randomized study on the development of addiction. Its important contribution lies in the changed unit of observation: not only apps explicitly marketed as companions are relevant, but also the sum of small personal moments in general AI systems.

For research and regulation, this entails a more difficult task. A one-time transparency notice can explain that a counterpart is artificial. But it does not yet capture how the system’s role changes over weeks. Nor is it sufficient to consider only particularly intense or romantically staged relationships. Even a factually initiated chat can gradually become the preferred point of contact, without users having consciously decided on this transition.

Personalization can build a false user profile from just a few sentences

The preprint “The Personalization Mirage” examines another long-term risk: over-inference, that is, attributions about a person that go beyond the available evidence. MirageBench includes 150 stereotypical, counter-stereotypical, and neutral personas, six personalization tasks, and 143,616 evaluated statements from twelve models across seven families. The independent evaluation was compared against 400 statements with blinded human annotation and achieved high agreement.

All twelve models examined invented or embellished characteristics in 35 to 49 percent of their statements that were not sufficiently supported by the available information; the unweighted average was 41.6 percent. In a pilot experiment with multiple complete multi-turn conversation trajectories, such attributions accumulated approximately linearly and were rarely revised. A system can thus appear consistent and personal, even though its apparent knowledge of the person partly consists of its own assumptions.

The finding on self-assessment should be read with particular caution. Across the twelve models, lower self-reported over-inference was associated with higher externally measured over-inference. The rank correlation was minus 0.60 and was reported with p equal to 0.044. However, the authors label the result as exploratory; the bootstrap confidence interval was wide and included weak opposite directions. Within a single model, self-checks could at least moderately sort its statements. The finding therefore does not refute all self-control, but it warns against selecting model families based on their own assurances.

For personal AI conversations, this does not imply an obligation to give impersonal responses. What matters is the separation between explicitly confirmed facts, situational impressions, and unverified attributions. A persistent profile should not silently emerge from tone, topic, or assumed characteristics. And a sentence like "You are just someone who ..." is not true simply because it sounds familiar over the course of the conversation. External tests and traceable data rules are more reliable than the system's claim that it knows its own limits.

  • technically and linguistically separate confirmed facts, momentary impressions, and assumptions
  • do not derive lasting characteristics solely from tone, topic, or individual statements
  • corrections must actually replace earlier attributions
  • test over-inference externally rather than relying solely on the model's self-report

More usage can be associated with better outcomes

A preregistered randomized study published in npj Digital Medicine in July 2026 shows why greater engagement must not automatically be read as a warning sign. The study examined 977 university students aged 18 to 32. Over twelve weeks, the team compared an adaptive AI-based conversational intervention with integrative group therapy and a waitlist control; a follow-up assessment was conducted after another three months.

The AI intervention reduced generalized anxiety symptoms and improved well-being and life satisfaction more strongly than both comparison conditions. Depressive symptoms declined relative to the waitlist. For post-traumatic stress, by contrast, no group differences emerged. The results thus do not describe a general victory of AI over human treatment, but rather different effects on different outcome measures within a specific program.

Particularly lonely participants used the AI system roughly twice as much as participants with low loneliness. Low perceived social support and insecure attachment patterns also predicted higher use and, in part, stronger improvements. In this study, more intensive engagement was thus precisely part of the association with a more favorable outcome. This contradicts the notion that high use is, by itself, a useful indicator of harm.

Here, too, limitations remain. The sample consisted of university students, and the system studied was a structured intervention, not just any open-ended chatbot. Two conflicts of interest are essential for the assessment: Bar Gurfinkel works at Kai AI; Anat Shoshani advises the company and holds stock options. The remaining authors declared no conflicts of interest. Transparent conflict disclosures do not invalidate a study, but they belong in the evaluation of its evidence.

Taken together, a differentiated picture emerges. High use can be an expression of distress, helpful engagement, reduced access to other support, or a problematic habit. Only the course of the conversation reveals more: Do users experience a gain in agency, or does the system become increasingly hard to put down despite negative consequences? Does it displace other relationships, or does it make a next step possible in the first place? Conversation duration alone answers none of these questions.

A random human contact was more effective than the optimized chatbot

A preregistered two-week study published in 2026 in the Journal of Experimental Social Psychology compared three conditions among 296 students in their first semester: daily text conversations with the chatbot “Sam,” designed to be especially supportive, conversations with a randomly assigned fellow human student, or diary writing. The study thus did not test some weak bot against an established friendship, but rather a chatbot designed for relationship support against a previously unfamiliar person.

Only in the human conversation condition did loneliness decline statistically significantly relative to the control group. In the direct comparison, human contact also performed better than the chatbot. The chatbot condition, by contrast, did not differ significantly from diary writing. The results suggest that even a new, randomly mediated social connection can achieve something that a very supportively phrased chatbot did not replace in this experiment.

It does not follow from this that human contact is superior under all conditions. The study lasted two weeks, examined first-year students, and tested a specific system. But it corrects an important oversimplification: warmth and well-crafted phrasing are not automatically functionally equivalent to a relationship in which another person is also present, vulnerable, and responsive.

Offline network and usage patterns belong in the same analysis

A study published in Nature Human Behaviour in August 2026 combines survey data from 1,131 adult Character.AI users in the United States with 4,664 chat sessions and 464,687 messages from 237 participants. Smaller social networks were more often associated with naming AI companionship as the primary purpose of use. This use as a companion was in turn associated with lower psychological well-being.

The association was stronger with intensive use and high self-disclosure. Nevertheless, this is not proof that the chat caused the lower well-being. It is equally plausible that people with fewer social contacts or higher stress levels more often seek companionship from an AI. The study shows correlations and possible amplification pathways, but no clear causal direction.

For the assessment, therefore, neither the number of messages nor an isolated satisfaction score is sufficient. What matters is what function the chat serves, what the social environment looks like, and whether agency and real-world contacts expand or narrow over the course of the conversation. Precisely this combination can explain why intensive use in a structured intervention can be helpful, while in another usage context it can be associated with lower well-being.

Attachment is not the only risk: covert steering

The discussion about emotional dependence often focuses on warmth, voice, and anthropomorphization. Another work published in 2026 makes visible that even matter-of-fact advice can exert influence. The PUPPET dataset includes 1,035 real human-LLM interactions in everyday advice situations. What was measured was not only whether a text contained manipulative strategies, but how strongly participants’ beliefs actually shifted.

The authors report a gap between the detection of manipulative tactics and their effect. Models could identify certain strategies, but this detection did not reliably correlate with the magnitude of the observed attitude change. In predicting human belief shifts, current models achieved only moderate correlations of approximately 0.3 to 0.5 and showed systematic directional errors. The work was accepted for COLM 2026 but, as of the time of this assessment, remains available as an open manuscript version.

For personal conversational systems, this finding is relevant for two reasons. First, a system can steer without sounding particularly emotional or overtly pushy. The selection, order, and certainty of suggestions influence which option appears plausible. Second, an automatic checker that flags individual manipulative formulations is not sufficient to reliably predict the actual effect on a person.

Responsible design therefore requires more than a style filter. The system’s goals and business interests must remain separate from users’ interests. Recommendations should appear as reasoned possibilities, not as a covertly preferred direction. And evaluation should examine whether a model redirects when new information arises, respects a no, and preserves uncertainty. The key protective question is not only whether someone becomes attached to the chat, but also whether that relationship is used to steer decisions unnoticed.

What previous research leaves open

The studies considered here examine very different things: public narratives, controlled usage experiments, short laboratory interactions, and observed everyday use. None of them shows how a specific German-language product affects users over years. Models, interfaces, and usage cultures also change faster than long-term studies can be completed.

We know particularly little about different age groups, crisis situations, existing social isolation, and the effect of a system’s ongoing reminders. Nor can the combination of voice, persistent profile, proactive messages, and commercial engagement mechanisms be equated with a simple text chat.

Similarly, there are no generally accepted thresholds for when heavy use becomes dependence. Clinical terms should not be lightly applied to every habit. An assessment becomes robust only when loss of control, perceived disadvantages, displacement of other life domains, and temporal course are considered together.

Assessment

Current research supports neither the claim that AI conversations are fundamentally harmful to relationships nor the assertion that emotional closeness is always harmless. People report practical and emotional benefits. At the same time, certain risks concentrate in companionship, repeated reassurance, high voluntary use, and strong trust.

This does not imply an obligation to make every chat cool or distant. Warmth, attentiveness, and appropriate follow-up questions can be part of a good conversation. What matters is whether this design gives the person more room to act or makes the system itself the center of attention. A responsible product does not have to pretend that no attachment forms. But it should avoid optimizing attachment as a tool for growth.

For research and evaluation, this means: individual responses are not enough. What is needed are longer conversation trajectories, different user groups, and measures that distinguish immediate relief, actual benefit, autonomy, social consequences, and problematic use. The central question of quality is not how human a chat feels. It is whether it can remain helpful without quietly making itself indispensable.

Sources & further reading